Introduction

DeepSeek Harness (DSH) provides plugin-based extensibility. For users who need to record quickly or find typing inconvenient, voice input is an immediate need. The dsh-voice-input plugin aims to address problems common in existing solutions—such as inaccurate recognition, the need to bypass network restrictions, and privacy issues caused by uploading audio to servers—by providing a local, streamlined speech-to-text experience.

Plugin Overview

Name: dsh-voice-input
Maintainer: jinhuoooo
Positioning: DSH voice input plugin. Click the microphone once and speak; the text is automatically inserted into the input box. Local Whisper engine, designed for users with limited typing proficiency.
Core capabilities:
* Local recognition: Integrates a local Whisper model and forces output in Simplified Chinese.
* Anti-hallucination: No hallucinated or made-up text is generated in silent environments.
* Privacy protection: All processing is local; audio does not leave the device.
* China-friendly: Models are downloaded from ModelScope without needing to bypass network restrictions.
* High performance: Resident process with millisecond-level transcription.
* Optional cloud: Supports switching to cloud APIs (Groq/SiliconFlow).

Installation and Enablement

Prerequisites

  • DSH: Any available version.
  • Python: 3.10+. During installation, select “Add Python to PATH”.
  • Microphone: System microphone permissions are enabled.

Installation Steps

Run the following command in the terminal to install the plugin:

dsh plugin --profile web add github:你的用户名/dsh-voice-input

web is the default Profile. If using another Profile, replace it.

First-time Use

The first time you click the microphone icon, the plugin automatically checks and installs dependencies (such as faster-whisper, modelscope, etc.) and downloads the Whisper model (default small, about 500MB). The interface shows progress during installation; you can use it once it is complete.

Usage

  1. Click the microphone icon on the right side of the input box.
  2. While speaking, watch the volume bar and recording duration below the button.
  3. Recording automatically stops and transcribes after 60 seconds.
  4. Click the red square to stop and insert the recognized text; click the gray × to cancel.

Configuration and Notes

  • Memory usage: Uses the small model by default, occupying about 500MB of memory. If this is a concern, use a different model or a cloud API.
  • Encoding issue: On Windows, Python’s default GBK output may cause garbled characters.
  • Permissions: Ensure that system microphone permissions are enabled.

Summary

This plugin provides DSH users with a local, streamlined, and Simplified Chinese voice input solution, suitable for users with limited typing efficiency or who are privacy-conscious.

  • Project URL: https://github.com/jinhuoooo/dsh-voice-input
  • License: MIT