Introduction¶
In the DeepSeek Harness plugin ecosystem, voice input in the Web profile usually requires calling an external API or relying on cloud services. The dsh-hold-to-talk plugin provides a key-free, fully offline voice input solution. It reuses the “hold to talk” interaction model from WeChat Desktop: press and hold the mouse in the input field, speak, and release to write the text into the draft. Recognition runs entirely locally, and the audio does not leave the machine.
What It Is¶
This is a client plugin for the DeepSeek Harness Web composer that provides local voice input capabilities. It is maintained by wangzhanchao883. It calls the SenseVoice model through sherpa-onnx and completes real-time speech recognition in the browser.
Core Features¶
- Hold-to-talk interaction: Press and hold the mouse in the input field for ≥400 ms to start recording, then release to insert the text into the draft.
- Local recognition: Uses sherpa-onnx to run the SenseVoice model, without an API key. It runs offline, and the audio does not leave the machine.
- Real-time preview: During recording, an overlay displays the real-time transcription text. This is only a preview and is not written into the draft.
- Cancellation mechanism: Slide up or press the Esc key during recording to cancel.
- Auto-completion: If the mouse is held for more than 60 seconds, recording automatically completes and the text is inserted.
- Thread optimization: Decoding runs in a Web Worker thread, avoiding blocking of the main thread.
Installation and Configuration¶
Installation¶
Make sure DeepSeek Harness is installed and the Web profile is running (dsh web), then run the following command:
dsh plugin --profile web add dsh-hold-to-talk
After installation, restart dsh web and refresh the page (Ctrl+Shift+R).
Configuration¶
The plugin’s default configuration is located in the profile’s cordis.patch.yml file. Add a config field under the dsh-hold-to-talk entry to override the default values:
- id: dsh-hold-to-talk
name: dsh-hold-to-talk
config:
holdThresholdMs: 400 # hold-to-talk trigger threshold
cancelSlidePx: 60 # slide-up distance to cancel
interimIntervalMs: 1500 # real-time preview interval
maxWindowSec: 6 # preview decoding window
minHoldMs: 250 # minimum trigger duration (prevent false trigger)
maxHoldMs: 60000 # maximum auto-completion duration
autoSend: false # whether to auto-send (default is insert only)
language: auto # language mode
useItn: true # whether to enable inverse text normalization
numThreads: 2 # number of engine threads
Typical Usage¶
- Start recording: Press and hold the mouse in the DSH Web input field and keep it still for ≥400 ms. An overlay will appear, showing “Release to send · Slide up to cancel”.
- Speak and preview: While you are speaking, the overlay displays the transcription text in real time (preview). The text is not immediately written into the input field.
- Insert text: Release the mouse, and the complete text from the overlay will be inserted into the draft in the input field.
- Cancel recording: Slide up by >60 px during recording, or press the Esc key, to cancel the current recording.
- Auto-complete: If the mouse is held for more than 60 seconds, recording automatically completes and the text is inserted.
- Pre-fetch the model: The first time it is used, the model is downloaded automatically (about 230 MB). To download it in advance, run the following command in the plugin directory:
npm run model:fetch
Notes¶
- Environment requirements: DeepSeek Harness must be running under the Web profile (
dsh web). The host machine must have Node.js >=22.19 or >=24 installed. - Browser compatibility: A Chromium-based browser (Chrome/Edge) is required, and microphone permission must be allowed.
- Disk usage: The first run requires downloading the model file, which uses about 230 MB of disk space.
- Preview vs. draft: The “real-time transcription” shown in the overlay is only a preview and is not written into the draft. Only the complete decoded audio result after releasing the mouse is inserted.
- Resampling mechanism: Audio resampling always starts from sample 0, and incremental resampling is not supported.
Summary¶
dsh-hold-to-talk provides lightweight, privacy-friendly voice input capabilities for DSH Web through local ASR and Web Worker decoding. For scenarios that require local processing of sensitive data or offline work, this plugin provides a practical solution.
Plugin directory: https://www.skillhub.cn/plugins/wangzhanchao883/dsh-hold-to-talk
GitHub: https://github.com/wangzhanchao883/dsh-hold-to-talk