DeepSeek Harness (DSH) uses a plugin-based architecture. The Web GUI provides a basic conversational interface but lacks voice input capability. dsh-voice, as an official plugin package, adds this functionality to the Web GUI.
What It Is¶
This is a DSH plugin maintained by KotDath. It integrates a microphone button and real-time waveform visualization into the editor, supports three transcription engines, and does not rely on the Handy app.
Core Features¶
- UI Integration: Displays a microphone button next to the send button in the editor. Clicking it opens a recording bar (with waveform).
- Operation Logic: The recording bar provides three actions:
✕— Cancel recording⏹— Stop and insert text into the input box↑— Stop and directly send to the chat
- Transcription Engines: Three built-in modes:
transcribe.cpp(default): Based on the Handy engine, supports 67 models.whisper.cpp: Runs locally.api: Compatible with OpenAI-style API transcription endpoints.
- Long Recording Processing: Recordings longer than 22 seconds are automatically segmented and processed.
- Independence: No need to install the Handy app; the engine can be downloaded automatically or loaded from PATH.
Installation and Activation¶
- Ensure DeepSeek Harness Web GUI is installed (Node.js environment).
- Run the installation command:
dsh plugin --profile web add github:KotDath/dsh-voice
- Restart the Web process to load the plugin:
dsh web
Typical Usage¶
- Configure the Engine: Open the Web GUI settings page and go to the Voice tab. The plugin automatically checks the engine status. Click “Download engine” to download it (supports multiple platforms).
- Select a Model: Choose one from the built-in catalog of 67 models. If it is not downloaded, click Download.
- Record: Click the microphone button in the editor to start recording.
- Send/Insert: After recording, choose either to stop and send directly, or simply insert the text.
Notes¶
- Runtime Environment: Only applies to DeepSeek Harness Web GUI.
- Host Dependencies: The host must have
bash,curl(for downloading the engine), andffmpeg(for audio conversion) installed. - Browser Limitation: Microphone access requires HTTPS or localhost.
- Security: The OpenAI API Key is configured only on the host side, not displayed on the page, and passed via environment variables.
- Data Storage: Plugin data is stored in the
.engine/,.models/, and.tmp/folders under the process startup directory.