The core design philosophy of DeepSeek Harness (DSH) is “everything is a plugin.” When writing prompts or interacting with agents, voice input often helps preserve the continuity of thought, while traditional speech-to-text tools are cumbersome to configure and carry privacy risks. dsh-asr-voice is a plugin for this scenario. It leverages cloud ASR speech recognition and LLM prompt optimization to turn spoken language directly into clean text.

Core Capabilities

  • Hybrid recognition architecture: By default, it prioritizes the browser-native Web Speech API (free and requires no API key). If unavailable, it automatically falls back to the configured cloud ASR. The cloud path supports multiple providers at the same time (OpenAI, Groq, SiliconFlow, Xiaomi MiMo, Tongyi Qwen-ASR, etc.).
  • Three-step setup wizard: The settings page provides a guided configuration flow. It can reuse existing DSH LLM credentials, so there is no need to manually manage API keys, and it supports a “Test Connection” self-check.
  • Real-time voice conversation: After enabling real-time mode, long-press the microphone button (about 450 ms) to enter conversation state. It supports half-duplex mode: the Agent reads its response aloud, playback ends, the microphone is automatically released, and long-pressing again interrupts it.
  • Privacy self-guarding: The API key is stored only in the DSH credential service; it is not written to plugin configuration files or the browser DOM. Audio is converted locally and then uploaded.

Installation and Configuration

Install using the official CLI:

dsh plugin --profile <profile> add <本插件路径或 GitHub 仓库>

After installation, find the “Voice Input” card in Plugin Configuration under DSH settings. The three-step configuration process is as follows:

  1. Select the recognition method: The default is “Cloud.” If you only need browser-based recognition, switch to “Browser Only.”
  2. Select the provider: Click a provider chip (such as OpenAI), and the BaseURL, model, and calling channel are automatically filled in.
  3. Validate the key: Check the key status. If it shows that DSH credentials are already being reused, no input is required. If not configured, enter the key and save it. The key is written only to DSH credentials and is not echoed back in the UI. Click “Test Connection” to verify availability.

Usage

  • Voice input: Click the microphone button on the right side of the input toolbar to start recording, then click it again to stop and recognize the audio.
  • Shortcut: The default shortcut is Ctrl+Shift+Space (Cmd-compatible on macOS) and can be configured.
  • Automatic submission after recognition: Enable this option in the behavior settings to achieve a Push-to-Talk-like experience.
  • LLM optimization: After recording stops, the cleaned text is immediately filled into the draft. LLM optimization runs silently in the background; once completed, it replaces the original text automatically without overwriting content the user has already edited.

Technical Details

  • Dependencies: At runtime, it only depends on the official @deepseek-ai/* peer packages, and it can be used standalone or together with other plugins.
  • Security: Audio files are converted locally before upload. The browser communicates with the plugin only through the private JSON proxy at /api/asr-voice/*.
  • Configuration: The plugin uses the dsh-asr-voice namespace and covers settings such as the recognition engine, optimization mode, real-time conversation toggle, and shortcut.

Conclusion

This plugin provides DSH users with a complete closed loop from voice input to text polishing. It is suitable for scenarios that require frequent interaction with agents, high input efficiency, or strong privacy protection.