Introduction¶
The DeepSeek Harness plugin ecosystem allows core capabilities to be extended through plugins. The dsh-realtime-voice plugin provides a lightweight way to connect browser audio input with LLM conversation capabilities, without deploying local models or a Python environment. It mainly solves the real-time flow from voice input to voice output, supports both Alibaba Cloud Bailian and OpenAI routes, and is suitable for scenarios requiring high concurrency and low-latency voice interaction.
Core Features¶
The plugin captures PCM audio in the browser, and the Host proxies ASR (voice recognition) and TTS (text-to-speech synthesis) services through a same-origin WebSocket. The specific capabilities are as follows:
- Dual-route support:
- Domestic route:
qwen3-asr-flash-realtime→ DeepSeek Harness →qwen3-tts-flash-realtime(Alibaba Cloud Bailian). - Global route:
gpt-realtime-2.1(OpenAI Realtime).
- Domestic route:
- Audio processing: TTS uses 24kHz PCM streaming playback. VAD (voice activity detection) has a default threshold of
0.85, and trailing silence detection is700ms. - Candidate interruption mechanism: The plugin rechecks candidate interruptions for text validity and similarity to the currently broadcast audio; false triggers restore the original broadcast.
Installation¶
It is recommended to install the prebuilt version directly from the version-tagged GitHub repository:
dsh plugin --profile web add github:zfu691531-hash/dsh-realtime-voice#v0.12.0
Uninstall command:
dsh plugin --profile web remove dsh-realtime-voice
Configuration and Usage¶
Provide the API key for the selected route in the DeepSeek Harness credential configuration:
* Qwen: DASHSCOPE_API_KEY.
* OpenAI: OPENAI_API_KEY.
* Voiceprint (optional): Tencent Cloud TENCENT_SECRET_ID and TENCENT_SECRET_KEY.
After installation, expand “Realtime Voice (Qwen / GPT)” in Harness’s “Settings → Plugins” to use it.
Status Check¶
Run the Doctor command to check the plugin status:
await fetch('/dsh-realtime-voice/status').then(response => response.json())
Notes¶
- Desktop limitation: The Electron shell of DeepSeek Harness Desktop rc.5 denies microphone permissions for the renderer process. When clicking the microphone button in the desktop app, open the same page in the default external browser (such as Chrome/Edge) and grant microphone access; the Host, session, and plugin still share the same instance.
- Security boundaries:
- API keys are read on a per-request basis only by the Host credential service and are not written to the browser,
localStorage, or logs. - HTTP and WebSocket routes only accept loopback and same-origin requests.
- Voiceprint is disabled by default and can only suppress automatic submission caused by background speech; it cannot defend against recorded replay, synthetic voice, or targeted impersonation.
- API keys are read on a per-request basis only by the Host credential service and are not written to the browser,
- Scope of use: The plugin runs under the permissions of the DeepSeek Harness process. Before installing, review the source code and license to ensure compliance with local security requirements.
Related Links¶
- Plugin directory: https://www.skillhub.cn/plugins/zfu691531-hash/dsh-realtime-voice
- Source repository: https://github.com/zfu691531-hash/dsh-realtime-voice