Introduction

DeepSeek Harness (DSH) uses a plugin-based architecture. If developers want to integrate real-time voice conversation capabilities, they typically need to handle low-level logic themselves, such as WebSocket upgrades, streaming speech recognition and synthesis, and microphone permission management. The dsh-voice-live plugin aims to solve this pain point by bringing together microphone capture, Volcengine streaming speech recognition (ASR), text-to-speech synthesis (TTS), and browser-side voice control to provide ready-to-use real-time full-duplex voice support.

What Is This

This is a plugin that provides real-time full-duplex voice capabilities for DeepSeek Harness, maintained by user tangzheng202202. By combining a /voice WebSocket route registered on the host with client-side rendering controls, it automates the complete flow from user speech to assistant response.

Core Features

The plugin includes the following core capabilities:

  • Real-time voice interaction: Start audio capture by clicking the microphone, and the assistant responds immediately and reads the response aloud.
  • Streaming processing: Integrates Volcengine streaming ASR and TTS with low latency.
  • Server-side endpoint detection: The server detects about 1.5 seconds of silence and automatically submits the recognition result, without requiring a manual stop.
  • Interruption and wake word: Supports interruption during playback and provides wake-word support (disabled by default).
  • Real-time subtitles: Renders recognized content in real time into the editor draft.
  • Browser echo cancellation: Uses native browser AEC technology to reduce echo.

Installation and Activation

This plugin has not yet been officially published via npm. It is recommended to build and run it directly within the DSH repository monorepo. Please clone the following fork repository and follow the DSH build process:

# 克隆 DSH Monorepo Fork
git clone https://github.com/tangzheng202202/deepseek-harness.git
cd deepseek-harness
# 在 monorepo 中构建并运行
# 具体路径需参考 DSH 官方文档,或自行 vendor 缺失的依赖包

Typical Usage

  1. Start voice: Click the microphone button in the editor; the system starts capturing audio and displays the recognition result in real time in the draft.
  2. Automatic submission: After you finish speaking, the system waits for about 1.5 seconds of silence and automatically submits the recognition result, without requiring a second click to stop.
  3. Reply-first: “Reply-first” mode is enabled by default; the system first reads a confirmation phrase aloud (such as “Okay, I’ll check that for you”), then plays the actual answer.
  4. Interruption interaction: While the assistant is replying, if the user speaks again, the system immediately stops playback and cancels the current in-progress TTS task, then starts new recognition.

Use Cases and Notes

  • Use cases: Suitable for developers who need to quickly integrate real-time voice conversation capabilities on the Web.
  • Notes:
    • Release status: The npm release is currently blocked upstream due to the missing @deepseek-ai/dsh-compact dependency, so it cannot be installed as a standalone plugin.
    • Wake word limitations: The wake-word feature depends on the Web Speech API and is disabled by default. Chrome cannot use Google cloud recognition services in mainland China, so only Safari (macOS) can be used; the local offline engine sherpa-onnx is currently unavailable.
    • Echo cancellation: Depends on native browser AEC and is limited by certain Chrome bugs, so the current mode is half-duplex (capture stops while playing).
    • Real-time subtitles: Currently shown only in the editor draft; floating subtitles are not available yet.
    • Server-side endpoint: Ensure that you use bigmodel_async and enable the enable_nonstream parameter.
  • The plugin runs with host process privileges. Please carefully review the source code and the MIT license before use.

Conclusion

dsh-voice-live provides DSH with a basic voice interaction skeleton and implements streaming voice capabilities through Volcengine. Developers can refer to its source code implementation and integrate it within a monorepo.