In the DSH ecosystem, voice input is an effective way to improve development and interaction efficiency. The current official DSH voice input feature is relatively basic and lacks practical features such as hotkeys, streaming display, and hot-word correction. dsh-voice-scribe is a third-party plugin maintained by PensiveFei, designed to provide a more complete local offline speech recognition experience, with support for cloud fallback and AI polishing.
Core Features¶
This plugin primarily addresses the convenience and accuracy of voice input, with the following core capabilities:
- Recognition engine: Local offline recognition by default (SenseVoice: zero configuration, no API key required, audio does not leave the local machine), automatic fallback to browser Web Speech, and optional cloud ASR (OpenAI-compatible).
- Interaction methods: Supports Alt hotkey (custom key combinations can be used), push-to-talk mode, and a 🎤 button.
- Real-time feedback: Provides real-time intermediate results (streaming) and recording level indication.
- Accessibility: Supports microphone device selection and diagnostics, canceling recording when switching windows, and multilingual support (Chinese/English/Cantonese/Japanese/Korean).
- Advanced features: Supports hot-word replacement tables (
hot.txt) and custom polishing prompts (reusing DSH models).
Installation and Activation¶
Before installing, make sure your DSH version is 0.1.0-rc.6 or higher.
dsh plugin --profile web add dsh-voice-scribe
After installation, restart the Web version of DSH to activate the plugin. This plugin package is compatible with both the Electron desktop version and the Web version.
Typical Usage¶
1. Basic Recording¶
The default hotkey is Alt. In the input box, press and hold or click and hold Alt to speak. Release it or press it again to append the recognition result to the end of the draft.
2. Custom Hotkey and Mode¶
Go to Settings → Voice Input:
- Hotkey: Click “Custom” to record any key combination (for example: Ctrl+Shift+V, F9, Alt+Space).
- Trigger method: Select “push-to-talk,” hold the hotkey to record, and release it to trigger automatic transcription.
3. Microphone Device Diagnostics¶
If the recognition result is empty, first check the “Microphone Device” in Settings. The default device for getUserMedia is determined by the browser (it may prioritize a muted virtual device), while the plugin supports specifying an exact device. The status bar during recording will show the device name that is actually in effect. If you see the prompt “input level ≈ 0,” it usually indicates that the capture side is muted (for example, a virtual sound card is occupying the input).
4. Hot-Word Correction¶
Edit $DSH_HOME/voice/hot.txt, with one rule per line:
# literal replacement (case-insensitive)
DeepSeek=deep seek|迪普西克
# regex replacement
/deep\s*seek/DeepSeek/gi
5. AI Polishing¶
Enable polishing in Settings and select the DSH model to reuse. The plugin first performs local rule-based pre-polishing (removing filler phrases), and then passes the text to the LLM for processing.
Recognition Engine Notes¶
- Automatic (default): Prefers local SenseVoice. If it is unavailable, automatically falls back to browser Web Speech.
- Browser Web Speech: Zero configuration, but depends on Google/Microsoft services and may be unavailable on networks in mainland China.
- Cloud ASR: Supports configuring multiple OpenAI-compatible endpoints (such as Groq, SiliconFlow, and Alibaba Cloud Bailian), with automatic failover on failure.
Known Limitations¶
- Input boxes containing @ references: When the input box contains chip references (such as
@file), the plugin writes the content by replacing the entire segment, which expands chips into plain text. It is recommended to send the current content or clear the draft before using dictation. - Mainland China network environments: Browser Web Speech depends on external services. In mainland China, local offline or cloud engines are usually required.
- Unofficial plugin: This plugin is not affiliated with DeepSeek company. Please inspect the source code and license yourself.
Summary¶
By reusing the DSH model system and the SenseVoice engine, dsh-voice-scribe provides developers with a voice input solution that is simple to configure, privacy-friendly (local audio does not leave the machine), and feature-rich. Compared with the official experimental voice input, it adds hotkeys, streaming display, hot-word correction, and AI polishing, making it suitable for developers who frequently use voice interaction.