Introduction¶
DeepSeek Harness (DSH) provides plugin-based extensibility. For users who need to record quickly or find typing inconvenient, voice input is an immediate need. The dsh-voice-input plugin aims to address problems common in existing solutions—such as inaccurate recognition, the need to bypass network restrictions, and privacy issues caused by uploading audio to servers—by providing a local, streamlined speech-to-text experience.
Plugin Overview¶
Name: dsh-voice-input
Maintainer: jinhuoooo
Positioning: DSH voice input plugin. Click the microphone once and speak; the text is automatically inserted into the input box. Local Whisper engine, designed for users with limited typing proficiency.
Core capabilities:
* Local recognition: Integrates a local Whisper model and forces output in Simplified Chinese.
* Anti-hallucination: No hallucinated or made-up text is generated in silent environments.
* Privacy protection: All processing is local; audio does not leave the device.
* China-friendly: Models are downloaded from ModelScope without needing to bypass network restrictions.
* High performance: Resident process with millisecond-level transcription.
* Optional cloud: Supports switching to cloud APIs (Groq/SiliconFlow).
Installation and Enablement¶
Prerequisites¶
- DSH: Any available version.
- Python: 3.10+. During installation, select “Add Python to PATH”.
- Microphone: System microphone permissions are enabled.
Installation Steps¶
Run the following command in the terminal to install the plugin:
dsh plugin --profile web add github:你的用户名/dsh-voice-input
webis the default Profile. If using another Profile, replace it.
First-time Use¶
The first time you click the microphone icon, the plugin automatically checks and installs dependencies (such as faster-whisper, modelscope, etc.) and downloads the Whisper model (default small, about 500MB). The interface shows progress during installation; you can use it once it is complete.
Usage¶
- Click the microphone icon on the right side of the input box.
- While speaking, watch the volume bar and recording duration below the button.
- Recording automatically stops and transcribes after 60 seconds.
- Click the red square to stop and insert the recognized text; click the gray × to cancel.
Configuration and Notes¶
- Memory usage: Uses the small model by default, occupying about 500MB of memory. If this is a concern, use a different model or a cloud API.
- Encoding issue: On Windows, Python’s default GBK output may cause garbled characters.
- Permissions: Ensure that system microphone permissions are enabled.
Summary¶
This plugin provides DSH users with a local, streamlined, and Simplified Chinese voice input solution, suitable for users with limited typing efficiency or who are privacy-conscious.
- Project URL: https://github.com/jinhuoooo/dsh-voice-input
- License: MIT