Preface¶
DeepSeek Harness (DSH) adopts a plugin-based architecture that allows functionality to be extended through plugins. When building agent applications, voice input is a high-frequency need. The following introduces the dsh-voice-input plugin, which implements speech-to-text using native browser capabilities and requires no additional server-side deployment.
Plugin Overview¶
- Name:
dsh-voice-input - Maintainer:
difimim - License: MIT
- Core Value: Adds a microphone button to the input box and uses the browser Web Speech API to convert speech into text in real time and fill it into the input box, enabling zero-cost voice interaction.
Installation and Activation¶
This is a standard DSH plugin bundle. Install it using the official CLI. After installation, the DSH process must be restarted for the change to take effect.
Installation command:
dsh plugin --profile web add @difimim/dsh-voice-input
After installation is complete, restart the DSH service:
dsh web
If you use another profile (for example,
tui), the command syntax is the same:
bash dsh plugin --profile tui add @difimim/dsh-voice-input
Core Features¶
- Input Box Toolbar Integration: Adds a microphone button to the left side of the input box toolbar, used to start or stop audio capture.
- Real-Time Recognition: Intermediate results are written into the input box draft while you speak, without waiting for recognition to complete.
- Append-Only Writing: Existing text is not cleared. Recognition results are appended to the end of the current draft, allowing them to coexist with manually typed content.
- Zero Cost and Zero Server: Reuses the built-in browser Web Speech API directly, without requiring self-hosted recognition services.
- Theme Adaptation: Button colors use DSH theme variables and automatically adapt to light and dark modes.
Usage¶
- Start Audio Capture: Click the microphone icon on the left side of the input box toolbar. The button turns red and displays a pulse animation.
- Speak and Input: Speak directly. Recognized text appears in real time in the input box.
- Stop Audio Capture: Click the microphone icon again (the icon changes to a square) to stop audio capture.
- Send: After confirming that the text is correct, press Enter to send as usual.
Technical Details and Notes¶
- Recognition Engine: Depends on the browser built-in
SpeechRecognition(orwebkitSpeechRecognition) interface. - Recognition Language: The default language is Chinese (
zh-CN). - Writing Method: Writes the draft using the
inputActions.setDraft()method provided by the input box Slot.
Compatibility and Limitations:
- Browser Requirements: Supported by Chrome / Edge / Safari (relatively recent versions); Firefox does not support the Web Speech API.
- Network Dependency: The recognition service requires an internet connection, and audio is sent to the browser vendor’s recognition server.
- Permission Prompt: The first time you click the microphone, a system authorization window appears, and you need to click “Allow”.
- Privacy Note: To achieve zero deployment, the current solution sends voice data to the browser vendor’s cloud recognition service. If you have extremely high privacy requirements or need fully offline capability, the current version does not meet those needs.
Summary¶
dsh-voice-input provides DSH with a lightweight voice input solution, suitable for developers who want to quickly add voice interaction capabilities to an existing DSH environment. For more details and source code, visit its directory page or GitHub repository.