DeepSeek Harness (DSH) uses a plugin-based architecture, allowing developers to extend functionality as needed. For developers who require natural interaction or want to quickly obtain summaries, voice input and output are common needs. The @allmodels/dsh-speech plugin provides a solution for this, supporting real-time voice-to-text conversion and converting the agent’s responses into voice summaries.
Core Features¶
The plugin provides the following core capabilities:
- Real-time voice input and transcription
- Voice answer summaries
- Concise voice summaries
- Microphone selection
- Speech recognition settings
- Speech summary settings
Installation and Enablement¶
- Run the installation command in the terminal:
dsh plugin --profile web add @allmodels/dsh-speech
- Restart DeepSeek Harness.
- Open Settings → Speech for configuration.
Typical Usage¶
- Voice Input: Click the microphone icon next to the send button to start speaking. The transcription is displayed in real time, and you can edit it before sending.
- Voice Summary: After the agent completes a task, the plugin converts the final message into a concise voice summary. You can enable auto-play in the settings or manually play it at any time.
Connecting to AllModels¶
Using this plugin requires connecting to the AllModels.io service. You can complete authentication in the following ways:
- Enter an email address, and AllModels will send a verification code.
- Use an existing AllModels API Key.
New users typically receive a certain amount of free credits.
Use Cases and Notes¶
- Use Cases: Suitable for developers who need to interact using voice during development or who want to quickly listen to summaries of the agent’s results.
- Environment Requirements:
- DeepSeek Harness web profile
0.1.5-rc.2or later. - Node.js
^22.19.0 || >=24.0.0.
- DeepSeek Harness web profile
- Permissions and Security: The plugin runs with the permissions of the current
dshprocess. It is recommended to review the source code and license (MIT) before installation.
Summary¶
This plugin integrates voice interaction into DSH workflows, simplifying the process of voice input and listening to summaries. For more details, please refer to the GitHub repository.