DeepSeek Harness (DSH) extends system capabilities through plugins. When developing or using DSH, you may need to quickly enter text in non-DSH environments, such as browsers, WeChat, or Notepad. The dsh-voice-input plugin addresses this need by using a local Whisper model to perform global voice transcription, and pastes the input directly at the current cursor location.

Plugin Overview

This is a DSH plugin maintained by lougyang. After it is mounted, it automatically starts the voice engine in the background. Triggered by a global hotkey, it can recognize speech and output text. The entire recognition process is completed locally, and data does not leave the local machine.

Installation and Enablement

Before installing, ensure that Python 3 is installed on your computer (3.10+ recommended).

  1. Run the following command to install the plugin:
   npx -p @deepseek-ai/dsh dsh plugin --profile web add github:lougyang/dsh-voice-input
  1. Restart the web profile. After the plugin is mounted, the voice engine starts automatically (a microphone icon appears in the system tray).

Core Features

The plugin mainly provides the following three capabilities:

  1. Speech Transcription
    Press and hold F2 to speak; after releasing, the text is automatically pasted at the current cursor location. Recognition is based on the local Whisper small model, and Chinese recognition works well.

  2. Audio Recording Only
    Press F3 to start recording and press it again to stop and save as a WAV file. This feature does not perform transcription; it only saves the audio.

  3. History and Settings
    Each transcription result or recording is saved to the history (including content, duration, and time). Users can view, copy, or clear records in the Settings menu accessed by right-clicking the tray icon.

Configuration Description

The plugin supports modifying parameters through the tray settings or configuration file. Changes saved in the tray settings take effect immediately. The following configuration items are supported:

Configuration Item Default Description
hotkey f2 Speech hotkey; supports combinations like ctrl+alt+space
mode hold Recognition mode: hold = press and hold to speak, toggle = press once to start / press again to stop
record_enabled true Whether to enable audio recording only
record_hotkey f3 Recording hotkey
record_dir empty Recording save directory; if empty, uses %USERPROFILE%\Recordings
model small Model size: tiny / base / small / medium / large-v3
language zh Language setting: zh / en / auto
device auto Device selection: auto (automatically uses GPU if an NVIDIA GPU is present) / cpu / cuda

Usage Example

In any application that supports text input (such as a browser, editor, or search box):

  1. Press and hold F2 to start speaking.
  2. Release F2.
  3. The system pastes the recognized text at the cursor position within 1~2 seconds.

Notes

  • Permission Requirements: DSH must be run with administrator privileges. Normal privileges cannot inject input into administrator windows, preventing text pasting in all applications.
  • Environment Dependencies: The first installation automatically creates a Python environment and downloads dependencies (approximately 2.5 GB).
  • Model Download: The first time F2 is pressed, the Whisper model is downloaded automatically (approximately 500 MB). In mainland China, the default uses the Tsinghua PyPI mirror and hf-mirror.
  • Performance: The first startup and model loading take time; subsequent recognition sessions typically take about 1~2 seconds.

References