Introduction¶
The core design philosophy of DeepSeek Harness (DSH) is “everything is a plugin.” When developing or debugging agents, textual responses may be information-rich but lack auditory feedback, making it difficult to intuitively evaluate tone or fluency. The dsh-tts plugin aims to solve this issue by providing text-to-speech (TTS) capabilities for the DSH Web UI, allowing agent responses to be read aloud in real time.
Plugin Positioning¶
goodandready/dsh-tts is a DSH plugin maintained by GooDAnDReaDY. Its core feature is to read out agent responses audibly in the Web UI. It adopts a multi-provider fallback chain mechanism, supporting multiple synthesis options from offline local engines to cloud APIs. The plugin itself is released under the MIT license.
Core Features¶
Local Offline Engines and Future Support¶
The plugin includes offline system engines, including Edge TTS, Piper, and eSpeak. These engines can run without cloud API keys, making them suitable for privacy-sensitive scenarios or environments with limited network access.
For neural models such as Kokoro-82M and F5-TTS, the plugin supports downloading weights and displaying download status, but it does not include a neural inference runtime. This means these providers cannot work directly at present; the plugin will report errors accurately and continue attempting the next provider in the fallback chain.
Real-Time Streaming Audio and Low Latency¶
The plugin achieves real-time streaming audio transmission with latency below 300 milliseconds. It uses SSE (Server-Sent Events) via the /dsh-tts/stream route to push the synthesized audio stream directly to the browser and uses AudioWorklet for PCM processing, ensuring playback without stutters or noise.
IT Terminology Dictionary and Multi-Agent Personas¶
The plugin includes a built-in IT terminology pronunciation dictionary, supporting the conversion of technical terms such as SQL, Nginx, and Kubernetes into standard spoken forms. At the same time, it supports assigning independent voice, speech rate, and SSML style to different subagents (Subagent), enabling multi-character dialogue.
Messenger Integration and Audio Export¶
The plugin supports integrating voice synthesis into the Messenger (Telegram/Discord) gateway. In addition, users can export an audio clip of any text via the exportAudioClip(text) function, downloading it directly from the history in the Web UI.
Installation and Enablement¶
Since the official documentation does not provide a specific CLI installation command, manual installation is recommended using one of the following methods:
- Install via directory or source code: Deploy the plugin code into the DSH plugin directory.
- Manual installation in Settings: In the DSH Web UI Settings interface, locate Plugin Management and manually specify the plugin source to install it. The Settings interface provides a real-time download progress bar and SHA-256 verification.
- Check dependencies: Before installation, ensure the following peer dependencies are installed:
@deepseek-ai/cordis@deepseek-ai/dsh-tools@deepseek-ai/dsh-credentials@deepseek-ai/dsh-host-webserver@deepseek-ai/dsh-settings@deepseek-ai/schemastery
Typical Usage¶
After enabling the plugin, configure the provider fallback chain in Settings. For example, you can prioritize the offline Piper engine and automatically switch to ElevenLabs or another cloud API when it is unavailable.
At the code level, you can convert text into an audio stream and export it by calling the exportAudioClip(text) method. If voice interaction needs to be enabled (such as VAD-triggered barge-in), the @goodandready/dsh-voice plugin must additionally be installed.
Notes¶
- Security: All API keys are processed only in the DSH backend and are not sent to the client browser, ensuring credential security.
- Neural model limitations: Kokoro-82M and F5-TTS currently only support weight downloading; the inference logic requires a separately configured runtime environment.
- Network dependencies: Using local engines such as Edge TTS often requires system-level CLI tools (e.g.,
edge-tts) or a Python environment; using cloud APIs requires network connectivity.
Summary¶
dsh-tts provides DeepSeek Harness with a complete TTS solution, balancing offline availability with high-performance cloud synthesis. Through provider fallback chains and real-time streaming, it can be seamlessly integrated into existing agent workflows and improve the interactive experience. For more details, see the GitHub repository or the plugin catalog.