Introduction¶
DeepSeek Harness is a plugin-based platform. When using the Web Profile for conversations or developing agents, voice input is a common requirement. This plugin provides localized speech-to-text capabilities for DSH, avoiding uploading audio streams to the cloud. It is suitable for scenarios with privacy or network environment requirements.
Plugin Overview¶
The plugin is named opensquad-ai/dsh-voice-input and is maintained by opensquad-ai. It is based on the SenseVoice model from FunAudioLLM and integrates a microphone button next to the conversation input box in the Web Profile. After recording, the plugin calls a local service to convert the audio into text and fills it into the input box.
Prerequisites¶
To use this plugin, the following environment requirements must be met:
- Python 3.10+: The plugin automatically installs dependencies such as
flask,onnxruntime,librosa,soundfile, andpyyaml, but Python itself must be installed manually. - Node.js >= 22.19.0: The version requirement for the DSH platform itself.
- Browser support: The browser must support the
MediaRecorderandgetUserMediainterfaces.
(Optional) ffmpeg: Used as a fallback for transcoding abnormal audio. The browser usually converts recordings into standard WAV, so ffmpeg is not required for normal use. However, it is recommended to install it to prevent compatibility issues in some browsers.
Installation and Enablement¶
Run the following command in the terminal to complete the installation:
dsh plugin --profile web add @opensquad/dsh-voice-input
After installation, restart DeepSeek Harness. The plugin will automatically detect the Python environment in the background, install any remaining dependencies, and start the backend services.
Core Features¶
This plugin mainly provides the following capabilities:
- One-click speech-to-text: A microphone button is displayed next to the input box. Click it to start recording, click it again to stop and transcribe, and the text is automatically filled into the input box.
- Automatic model management: On first use, the plugin automatically downloads the SenseVoice model (about 230 MB) and displays a circular progress indicator and real-time percentage on the button.
- Separate management button: In addition to the record button, the plugin provides a separate management button for viewing installed models or uninstalling and re-downloading them.
- Automatic service startup: The plugin automatically starts the local Flask service (port 7101) and the download gateway (port 7102), without manual configuration.
Typical Usage¶
After completing installation and restarting, use the plugin as follows:
- Initialization: After restarting dsh, wait a moment. The background will automatically install Python dependencies and start the services.
- First use: Click the microphone icon, confirm the model download, and wait for the progress bar to complete. At this point, the services will be ready automatically.
- Recording and transcription: Click the microphone again to start recording. After finishing speaking, click stop. The transcribed text will automatically appear in the input box.
Applicable Scenarios and Notes¶
- This plugin runs under the Web Profile of DeepSeek Harness.
- The plugin runs with the permissions of the current dsh process. It is recommended to check the source code and license before installation.
- Downloading the model and installing dependencies may be slow on first use. Please be patient.
Technical Details¶
The backend logic of the plugin is managed automatically by the plugin. The startup flow is as follows:
- Detect the Python environment.
- Automatically run
pip installto install dependencies. - Start the download gateway (
gateway.py:7102). - Check whether the model already exists; if not, download it and start the transcription service (
service.py:7101).
License¶
This plugin is open-sourced under the MIT License.