Introduction

DeepSeek Harness (DSH) adopts a plugin-based architecture designed to extend system capabilities. DSH Realtime Voice is an official real-time voice Agent plugin, providing full-duplex voice interaction and interruption capabilities. It supports starting, appending to, correcting, or stopping the work of a DSH Agent by voice.

Core Features

Interaction and Interface

After installation, the DSH native 3080 WebUI will display a dial button to the right of the send button, using a blue circular call icon with the same size and color scheme. The interface includes an independent, draggable voice floating window that supports expanding and collapsing, and displays a dynamic waveform and timer at runtime. Users can continue the conversation, interrupt the Agent’s playback, or ask for progress during a call.

Model Switching and Interruption

In the WebUI under “Settings → Plugins → DSH Realtime Voice,” you can switch between Qwen Audio Realtime Flash/Plus models with one click, effective starting from the next call. The plugin supports fast acoustic interruption (using a three-layer interruption mechanism comprising browser-local voice onset detection, explicit Host cancellation, and Bailian VAD) and intelligent semantic turn switching.

Security and Architecture

The plugin automatically detects DASHSCOPE_API_KEY and also supports secure writing via the official DSH write-only credentials API. The API key does not enter the browser bundle and is not written to cordis.patch.yml, browser code, or Git. The plugin does not modify the DSH source code, does not start separate background processes, and cleans up all resources upon uninstallation.

Semantic Handoff and Closed Loop

The plugin uses a “dual-plane” runtime: Qwen Audio Realtime handles low-latency listening, speaking, questions and answers, and VAD; the DSH Agent handles tools, projects, and execution. Only real work such as files and applications is handed to DSH through Function Call. The plugin subscribes to DSH’s approval/requested and question/requested events to implement a closed loop for approvals and follow-up questions; it also subscribes to turn/end and other events to implement closed loops for progress and final states.

Protocol and Audio

The plugin provides the dsh.voice.v1 (WebUI) and dsh.voice.direct.v1 (WeChat/native client) control protocols. Downstream PCM is reassembled per response into ordered 40 ms (1920-byte) frames at 24 kHz/mono/s16le.

Installation and Configuration

Install using the official CLI:

dsh plugin --profile web add github:martinbear1/dsh-realtime-voice#v0.1.0-alpha.9

The uninstall command is as follows:

dsh plugin --profile web remove @harness-remote/dsh-realtime-voice

Configure DASHSCOPE_API_KEY in the WebUI under “Settings → Plugins → DSH Realtime Voice.”

Usage Examples

  1. Click the blue dial button that appears next to the WebUI input box.
  2. Continue the conversation; the Agent will transcribe and read responses in real time.
  3. During the call, interrupt the Agent’s playback or ask it to perform a specific task.
  4. Switch to the Qwen Audio Realtime Flash or Plus model on the plugin settings page.

Notes

  • The plugin runs with the permissions of the current DSH process. Check the source code and license before installation.
  • The plugin does not modify the DSH source code, does not start separate background processes, and cleans up all resources upon uninstallation.
  • Plaintext keys cannot be read back from the browser side.
  • Mini Program V1 supports only foreground real-time calls.
  • Automated tests will not collect or upload ambient audio without authorization.

Summary

DSH Realtime Voice combines the low-latency interaction capabilities of Qwen Audio Realtime with the execution capabilities of the DSH Agent. Through semantic handoff and closed-loop mechanisms, it implements a voice-driven Agent workflow. Developers can quickly integrate it into the DSH environment using the commands above.