Introduction

DeepSeek Harness (DSH) uses a plugin-based architecture. In model inference scenarios, manual input has latency and cannot handle complex continuous voice commands. The dsh-voice plugin aims to solve this problem. It provides zero-latency streaming dictation and voice memo functionality for DSH Web UI, and supports multi-provider fallback chains to ensure stable and real-time voice input.

Plugin Overview

dsh-voice is an open-source plugin focused on voice interaction. It is maintained by GooDAnDReaREADY and licensed under the MIT License. It works through a browser client and a host-side component to perform speech-to-text conversion and injects the results into DSH’s conversation flow.

Core Features

The plugin mainly provides the following capabilities:

  • Streaming dictation and VAD: Uses Voice Activity Detection (VAD) to segment speech fragments at natural pauses, enabling real-time typing.
  • Voice memos and auto-send: Supports voice memo functionality. After recording ends, it enters a cancellation countdown (default 4 seconds). When the countdown ends, the message is sent automatically; it can be canceled within the countdown.
  • Press-and-hold to speak: Supports mouse press/release gestures or keyboard shortcuts (such as Ctrl) to press-and-hold to speak. Releasing or pressing Esc can send or cancel.
  • Zero-latency browser subtitles: Uses the Chrome Web Speech API to provide real-time floating subtitles locally in the browser, with no network latency.
  • Multi-provider fallback chain: Built-in fallback chains for providers such as Deepgram, Groq, HuggingFace, and whisper.cpp, automatically switching when the primary service is unavailable.
  • Contextual dictionary injection: Automatically extracts code variables and identifiers from the conversation draft to improve recognition accuracy for technical terms.
  • Embedded audio player: Provides audio preview, progress bar, and playback controls in chat and editors.
  • Hardware noise suppression toggle: Supports browser-level hardware noise suppression, echo cancellation, and automatic gain control.
  • Provider health monitoring: Provides a real-time dashboard for latency (milliseconds), success rate, and error messages.
  • Zero API leakage: API keys are resolved on the host side via ctx.credentials and are never sent to the browser client.
  • Offline local Whisper: Supports starting a local whisper.cpp server for offline transcription.
  • SenseVoice-ONNX / Sherpa-ONNX: Supports version 0.8.11 of the SenseVoice-ONNX and Sherpa-ONNX engines, providing non-autoregressive fast speech recognition.
  • Real-time audio streaming: Achieves low-latency audio streaming through WebSocket bridging.
  • Audio visualization: Supports audio visualization effects such as liquid waveforms and dynamic light spheres.

Installation and Enabling

The plugin is published via npm. After installation, it runs as part of the DSH ecosystem.

Typical Usage

  • Streaming dictation: Click the microphone icon and speak. The plugin automatically inputs speech segments into the editor at pauses.
  • Voice memo: Click the wave icon. After recording ends, wait for the countdown to finish, then the message is sent automatically.
  • Press-and-hold to speak: Press and hold the wave icon (mouse) or the Ctrl key (keyboard) to start recording. Release it or press Esc to cancel.

Applicable Scenarios and Notes

  • Language support: Without a translation plugin installed, the Web UI interface defaults to English.
  • Privacy and security: All API keys are resolved on the host side. The browser side only handles audio capture and subtitle display, meeting security requirements.
  • Client build: The browser client code is compiled and generated from lib/client-src/*.js fragments using the npm run build:client command.

Summary

dsh-voice uses VAD segmentation, multi-provider fallback, and localized deployment (such as whisper.cpp) to provide a complete voice input solution for DeepSeek Harness. Its zero-latency features and host-side key management mechanisms make it suitable for development scenarios with high requirements for real-time performance and security.

Plugin Directory
GitHub Repository