Introduction

DeepSeek Harness (DSH) uses a plugin-based architecture and supports extending features on the desktop side. For LLM replies, reading long texts or debugging can become a bottleneck. dsh-tts-flash is an LLM reply read-aloud plugin designed for the DSH desktop client. It uses a streaming output side channel to synthesize and play audio locally while the AI is generating its response.

Core Capabilities

The plugin provides the following core features:

  • Lossless streaming read-aloud: The AI streams output while audio is synthesized and played locally, with pause as needed. Supports sentence splitting and short-sentence merging for Chinese, English, and Japanese; long replies automatically catch up.
  • Multi-engine support: Includes Microsoft Edge TTS by default, and also supports OpenAI-compatible cloud TTS services.
  • Plug-and-play engines: Engines can be added manually in the settings panel or automatically loaded from a JSON declaration file.
  • Waiting-phase audio system: While the AI is thinking (waiting for generation), the floating bar rotates fun phrases every 2.6–3.5 seconds, providing synchronized text and audio feedback.
  • Audio settings panel: Read-aloud toggle, voice model, volume, speed, subtitle font size, and other settings take effect in real time.
  • Read-aloud floating bar: Draggable with position memory; supports flowing-light subtitles and single-pass scrolling subtitles.
  • Persistent settings: Runtime configuration is saved independently and not lost after restart or reinstallation.

Installation

Build the package with npm:

npm pack
# 生成文件:dsh-tts-flash-0.1.0.tgz

After obtaining the .tgz file, register it using DSH’s plugin installation flow.

Engine Configuration

The plugin supports two configuration methods for engines.

1. Edge TTS (Default)

No extra configuration is required; select a Microsoft voice from the dropdown to use it. This service requires network connectivity.

2. OpenAI-Compatible Cloud TTS

In the Cloud Engine section of the settings panel, enter baseURL, API Key, and model id to add it. The model ID must exactly match the provider documentation.

3. JSON Declaration File

Place a JSON declaration file in the ~/.dsh/tts-flash/engines/ directory. The plugin loads it automatically at startup.

Example content:

{
  "id": "my-engine",
  "label": "我的引擎",
  "kind": "openai",
  "url": "https://api.example.com/v1",
  "apiKey": "sk-...",
  "model": "tts-model-id",
  "stylePrompt": "可选:音色描述/风格指令"
}

Waiting-Phase Audio

When the AI has started generating a reply but audio has not been generated yet, the floating bar rotates through fun phrases and reads them aloud. Audio files are cached in the ~/.dsh/tts-flash/cache/thinking/ directory.

  • Phrase pool: Includes 5 built-in fun phrases.
  • Independent model: The waiting-phase audio can use an independent audio engine (default: Edge Xiaoyi).
  • Actions: Supports batch generation of audio files or clearing the cache.

UI and Settings

The plugin provides a visual settings panel and a floating bar.

Floating Bar

  • Function: A pure-text capsule that supports dragging and position memory.
  • Effect: Long sentences display single-pass scrolling subtitles, with scroll duration equal to the actual audio duration.

Settings Panel

  • Read-aloud control: Enable or disable read-aloud in real time.
  • Audio parameters: Volume (0%–200%) and speed (Edge server cap: +100%).
  • Subtitle style: Flowing-light effect and font size (10–30px).
  • Cloud engine management: Add engines, edit voice descriptions, or delete engines.
  • Waiting-phase audio management: Select independent engines and clean up files.

Notes

  • Autoplay blocking: Browsers may block audible autoplay. First interact with DSH (send a message) before the read-aloud feature can be triggered normally.
  • Edge TTS limitations: This is an unofficial Microsoft API, with no SLA guarantee and unavailable when offline.
  • Speed limitation: Edge TTS server-side speed is capped at +100%.
  • Permission limitations: The desktop web server only allows requests from the DSH rendering process; all external access is denied.

Summary

dsh-tts-flash converts DSH text output into real-time audio through streaming synthesis and multi-engine support. Its waiting-phase audio and floating bar improve interactivity and immersion. For more details, see the GitHub repository.