DeepSeek Harness (DSH) extends system capabilities through plugins. When developing or using DSH, you may need to quickly enter text in non-DSH environments, such as browsers, WeChat, or Notepad. The dsh-voice-input plugin addresses this need by using a local Whisper model to perform global voice transcription, and pastes the input directly at the current cursor location.
Plugin Overview¶
This is a DSH plugin maintained by lougyang. After it is mounted, it automatically starts the voice engine in the background. Triggered by a global hotkey, it can recognize speech and output text. The entire recognition process is completed locally, and data does not leave the local machine.
Installation and Enablement¶
Before installing, ensure that Python 3 is installed on your computer (3.10+ recommended).
- Run the following command to install the plugin:
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:lougyang/dsh-voice-input
- Restart the
webprofile. After the plugin is mounted, the voice engine starts automatically (a microphone icon appears in the system tray).
Core Features¶
The plugin mainly provides the following three capabilities:
-
Speech Transcription
Press and hold F2 to speak; after releasing, the text is automatically pasted at the current cursor location. Recognition is based on the local Whispersmallmodel, and Chinese recognition works well. -
Audio Recording Only
Press F3 to start recording and press it again to stop and save as a WAV file. This feature does not perform transcription; it only saves the audio. -
History and Settings
Each transcription result or recording is saved to the history (including content, duration, and time). Users can view, copy, or clear records in the Settings menu accessed by right-clicking the tray icon.
Configuration Description¶
The plugin supports modifying parameters through the tray settings or configuration file. Changes saved in the tray settings take effect immediately. The following configuration items are supported:
| Configuration Item | Default | Description |
|---|---|---|
hotkey |
f2 |
Speech hotkey; supports combinations like ctrl+alt+space |
mode |
hold |
Recognition mode: hold = press and hold to speak, toggle = press once to start / press again to stop |
record_enabled |
true |
Whether to enable audio recording only |
record_hotkey |
f3 |
Recording hotkey |
record_dir |
empty | Recording save directory; if empty, uses %USERPROFILE%\Recordings |
model |
small |
Model size: tiny / base / small / medium / large-v3 |
language |
zh |
Language setting: zh / en / auto |
device |
auto |
Device selection: auto (automatically uses GPU if an NVIDIA GPU is present) / cpu / cuda |
Usage Example¶
In any application that supports text input (such as a browser, editor, or search box):
- Press and hold F2 to start speaking.
- Release F2.
- The system pastes the recognized text at the cursor position within 1~2 seconds.
Notes¶
- Permission Requirements: DSH must be run with administrator privileges. Normal privileges cannot inject input into administrator windows, preventing text pasting in all applications.
- Environment Dependencies: The first installation automatically creates a Python environment and downloads dependencies (approximately 2.5 GB).
- Model Download: The first time F2 is pressed, the Whisper model is downloaded automatically (approximately 500 MB). In mainland China, the default uses the Tsinghua PyPI mirror and hf-mirror.
- Performance: The first startup and model loading take time; subsequent recognition sessions typically take about 1~2 seconds.