Introduction¶
DeepSeek Harness (DSH) uses a plugin-based architecture that allows community developers to extend its functionality. @motong/dsh-voice is a community plugin that adds voice capabilities to DSH, maintained by the user motongv. It does not depend on any external API key and aims to provide developers with convenient speech-to-text (STT) and text-to-speech (TTS) capabilities, filling the gap in DSH’s native voice interaction.
Core Features¶
The plugin mainly provides the following capabilities:
- Voice input: Adds a microphone button to the input toolbar, supporting click-to-speak or invocation via a shortcut key.
- Shortcut voice input: The default shortcut is
Ctrl+Shift+Space, and custom key combinations are supported. - Response reading: Uses Microsoft Edge neural voices to read AI responses, with voices that sound more natural than system defaults.
- Voice management: Supports switching voices and previewing sound in settings; voice changes take effect immediately.
Installation and Enablement¶
The plugin is published on npm. Use the following command to install it:
dsh plugin --profile web add @motong/dsh-voice
After installation, you need to restart DSH to load the dependencies (the dependency ws is installed automatically by npm).
Typical Usage¶
- Voice input (click): Open any conversation, click the microphone icon 🎤 to the left of the input box, and speak into the microphone (Chinese is recognized by default). Recording stops automatically when you finish, and the text is inserted into the input box.
- Voice input (shortcut): Press
Ctrl+Shift+Space(or another key combination configured in settings) to start or stop voice input. - Adjust and preview: Go to Settings → Voice. You can change the shortcut, select a different voice (default: “Xiaoxiao”), or preview the sound.
- Read responses: Below any completed AI response, click the play button ▶ to have the browser play the synthesized speech; click the stop button ■ to stop playback.
Applicable Scenarios and Notes¶
- Browser environment: The plugin relies on native browser speech recognition. Make sure you are using the Chrome or Edge browser and have granted microphone permissions.
- Network requirements: The reading feature requires an internet connection and calls Microsoft Edge voice services. If the network or proxy blocks
speech.platform.bing.com, the reading feature will fail. Ensure direct network access. - Privacy note: This plugin does not access or store any API keys. During reading, the response text is sent to Microsoft Edge voice services for audio synthesis, while the recognition process is handled by the browser vendor.
- Open-source license: The plugin is open-sourced under the MIT License.
Closing¶
By introducing native browser APIs and Microsoft Edge TTS services, @motong/dsh-voice provides lightweight voice enhancement for DSH. Following the steps above, you can integrate it into your development environment for a more natural interaction experience.