Preface¶
DeepSeek Harness (DSH) Desktop aims to extend the capabilities of intelligent agents through a plugin-based architecture. During development or use, developers often need to keep staring at the screen to confirm the model’s thinking process, execution status, or output results. dsh-voice-mini is a voice feedback plugin designed for this purpose. It enables the assistant to “speak” using TTS technology, rather than mechanically reading long text, helping users track task progress and decision points auditorily.
What Is This¶
dsh-voice-mini is a voice feedback plugin for DeepSeek Harness, maintained by CroissanTTs. It addresses the question of “how can an agent proactively inform users of its current status.” With its built-in speak tool and templated broadcast mechanism, the plugin supports features such as the model deciding what to say, status broadcasting, and a native macOS floating pet, and provides multiple TTS backends from cloud-based to local.
Core Features¶
The plugin implements voice interaction through the following mechanisms:
-
Three speech modes
- speak tool: The model autonomously decides what to say (e.g., results or progress), presented in natural language rather than reading out the full reply.
- readReplies: Reads the assistant’s messages word by word (disabled by default), and is mutually exclusive with the
speakmode. - Verbalizer: After a conversation turn ends, a lightweight LLM rewrites the reply into a single emotionally toned sentence (e.g., “Good news, all tests passed”), with streaming generation support.
-
Status broadcasting
- Provides templated announcements for events such as approvals, questions, turn start/end, and pending items.
- Includes dedicated phrases for abnormal endings (interruption, blocking, errors, etc.).
-
Session voice
- Each session is automatically assigned a unique voice based on an FNV-1a hash, with a ±6% rate jitter added, removing the need for manual voice selection.
-
Auditory assistance
- Chimes: System-notification-style alert sounds (e.g., macOS glass sound), played before the voice.
- Monitoring panel: Provides an SVG line chart view of token consumption, latency, and system statistics.
- Native macOS Floating Pet: A native macOS floating window that displays the current utterance and pending items.
-
TTS backends
- Supports four backends: edge (default, cloud-based, Chinese and English), kokoro (local, English only), say (macOS offline), and fake (testing).
Installation and Activation¶
The installation process involves building and configuring links:
- Build the plugin
Run the build command in the plugin directory:
npm run build
This generates `lib/client.js` and the required configuration files.
- Build the native pet (optional, macOS only)
If you need the native floating pet feature, you must build it with Swift:
DEVELOPER_DIR=/Library/Developer/CommandLineTools swift build --package-path pet -c release
- Link it to the DSH configuration
In DSH’spackage.json, add the dependency and specify the entry:- Add a
link:dependency independencies. - Specify
entryinbundleto point to./cordis.patch.yml. - Restart DSH Desktop.
- Add a
Typical Usage¶
The plugin uses presets and configuration to provide voice feedback for different scenarios:
-
Configure presets
Configure three presets via the settings panel or API:- Immediate: speech rate +18%, announces every step.
- Default: rate 0%, announces only when delivering results or encountering a block.
- Minimal: rate -5%, volume -30, announces only approvals and questions, while the assistant remains silent (chimes only).
-
Handling blocking events
When the model cannot continue execution (e.g., when user input is required), a templated announcement prompts the user, avoiding silence while the model waits. -
Switching language
The plugin supports full zh/en internationalization switching. UI text, broadcast phrases, and Verbalizer prompts are all controlled by the locale, and can be manually overridden in the settings panel instead of the auto-detected browser language. -
Verifying features
Use the test script to verify TTS and routing:
npm test
Use Cases and Considerations¶
This plugin is suitable for developers or users who need multitasking and rely on auditory monitoring to track task progress.
- Network dependency: Using the
edgebackend requires an internet connection;kokorouses local inference but supports English only. - Mode exclusivity: The
readRepliesandspeakmodes cannot be enabled at the same time. - System requirements: The native floating pet feature requires a macOS environment and successful compilation of the Swift code.
- Permissions and security: The plugin runs with DSH process privileges. Before installation, please review the source code and license (MIT).
Conclusion¶
By separating model decisions from status templates, dsh-voice-mini provides DeepSeek Harness with a flexible auditory feedback layer. Whether using cloud TTS or a local offline solution, it extends agent interaction from the screen to speech, improving the user experience.
Project links:
* GitHub
* Community directory