Introduction

The core philosophy of DeepSeek Harness (DSH) is “everything is a plugin.” When performing model inference or interactions on the Web, frequently switching input fields or clicking buttons can interrupt the workflow. dsh-voice-assistant is a plugin designed for DSH Web, intended to address the complexity of Web-based input. It is maintained by supersyh-sss and provides keyword wake-up, voice dictation, dictation-based editing, and Chinese text-to-speech (TTS) capabilities.

Core Features

The plugin mainly provides the following capabilities:

  • Keyword wake-up: Supports custom wake words (default is “Little Whale”) and includes tolerance for homophones and near-homophones.
  • Voice dictation: After waking up, you can dictate continuously, and sentences are automatically added to the input field one by one.
  • Dictation-based editing commands: Commands start with “Execute” (e.g., “Execute Send”) to avoid accidental triggers. Common words in the body text will not trigger commands.
  • Chinese TTS: After a model response is complete, it is automatically read aloud in Chinese, with interrupt support.
  • Local offline recognition: Uses the sherpa-onnx WASM engine by default, runs locally in the browser, and supports offline use.
  • Hot configuration: Allows immediate adjustments to wake words, recognition engine, and other options in the settings page without restarting.

Installation and Activation

Install the plugin globally with npm and register it to the DSH Web profile.

npm i -g dsh-voice-assistant

dsh plugin --profile web add github:supersyh-sss/dsh-voice-assistant

After installation, restart dsh web; the plugin loads automatically. Verify that dsh.profile.bundles in ~/.dsh/profiles/web/package.json includes dsh-voice-assistant.

Typical Usage

  • Wake up: Say “Little Whale” to the microphone.
  • Dictate: After waking up, speak your content; it is automatically filled into the input field.
  • Commands: Say “Execute Send” to send a message, or “Execute Clear” to clear the draft.
  • TTS: After a model response is complete, it is automatically read aloud in Chinese.

Technical Details

  • Recognition engine: Uses sherpa-onnx WASM for local inference by default; no proxy is required on networks inside China. If local resource loading fails, it automatically falls back to the browser-native Web Speech API.
  • Tolerance mechanism: Supports homophones (e.g., “Xiao Jing”) and near-homophones (e.g., “Xiu Jing”, which requires a second confirmation in strict mode).
  • Accidental trigger prevention: All commands must start with “Execute”. For example, “Send” is only entered as content, while “Execute Send” triggers the send action.

Notes

  • This plugin requires browser microphone permission.
  • The first local recognition use requires downloading approximately 38MB of resource files.
  • The default resource server uses the jsDelivr CDN. If access from mainland China is restricted, you can manually specify the resource address on the settings page.

Conclusion

This plugin is suitable for developers who need long-running conversations on the Web or want hands-free operation. For more details, refer to the plugin directory or GitHub repository.