Introduction

DeepSeek Harness (DSH) is a tool for debugging and developing agents. In agent development, voice input is a common requirement, but it usually relies on external APIs or complex local service deployment. dsh-voice is a plugin that introduces local, private speech-to-text capabilities to DSH, running without external binaries.

Plugin Introduction

dsh-voice is a single-surface plugin, intended to provide voice input for DeepSeek Harness in the Web UI and TUI.

  • Maintainer: wencharmwang
  • License: MIT
  • Core value: Provides a host-side, ONNX-based local Whisper service and integrates a microphone button in the browser.

Installation and Enabling

This plugin is added to a DSH Profile as a Bundle.

  1. Install from NPM:
    dsh plugin --profile web add dsh-voice
  1. Or install directly from GitHub:
    dsh plugin --profile web add github:wencharmwang/dsh-voice

Note: When installing from Git with pnpm >= 10, if the prepare script fails, configure allowBuilds in the Profile’s pnpm-workspace.yaml.

After installation, the plugin automatically enables the voice Row. You can also declare the Row manually in ~/.dsh/profiles/web/cordis.patch.yml.

Configuration

All configuration is optional and can be modified through a Patch file. Configuration items include model, language, data type, and cache directory.

- id: voice
  config:
    model: onnx-community/whisper-medium  # Hugging Face Whisper ONNX ID 或本地目录
    language: auto                       # 'auto' | 'zh' | 'en' | ...
    dtype: q8                            # 'q8' | 'fp32' | 'q4'
    dir: ''                              # 模型缓存及录音目录,留空则使用默认路径

Model-level precedence: config files take priority over Schema defaults.

Features and Usage

Host-Side Service (ctx.stt)

The plugin exposes the ctx.stt service on the host side, providing the following interfaces:

  • status(): Checks readiness state and current model without loading/downloading the model.
  • ensureModel(onProgress): Ensures the model is downloaded and loaded.
  • transcribe(input, options): Converts audio input to text. Supports Float32Array (16kHz mono), Buffer (PCM/WAV), or a file path.
  • startRecording() (TUI): Starts recording and returns the WAV file path along with stop/cancel methods.

Web UI Interaction

A microphone button is integrated into the Composer in the Web UI. Recording is completed locally in the browser, decoded into 16kHz mono PCM, and then sent to the host side via an HTTP request.

HTTP Routes

In Web mode, the following routes are exposed:

  • POST /voice/transcribe: Receives audio data and returns transcribed text. Supports the query parameter ?lang=<iso> to specify the language.
  • GET /voice/status: Checks the model loading status.

Model Support

The plugin supports multiple Whisper ONNX models downloaded from the Hugging Face Hub, cached in ~/.dsh/voice/models.

  • onnx-community/whisper-medium (~0.8 GB, default)
  • onnx-community/whisper-small (~250 MB)
  • onnx-community/whisper-large-v3 (~1.6 GB)

Notes

  • TUI recording dependency: Microphone recording in TUI mode requires system-level recording tools (such as ffmpeg, sox, or arecord).
  • Model cache: The model is downloaded automatically on first use and cached in the specified directory afterward.
  • Ecosystem positioning: This plugin is maintained by a community developer and has no direct affiliation with official DeepSeek. Please review the source code and license before use.