Preface

DeepSeek Harness (DSH) Desktop aims to extend the capabilities of intelligent agents through a plugin-based architecture. During development or use, developers often need to keep staring at the screen to confirm the model’s thinking process, execution status, or output results. dsh-voice-mini is a voice feedback plugin designed for this purpose. It enables the assistant to “speak” using TTS technology, rather than mechanically reading long text, helping users track task progress and decision points auditorily.

What Is This

dsh-voice-mini is a voice feedback plugin for DeepSeek Harness, maintained by CroissanTTs. It addresses the question of “how can an agent proactively inform users of its current status.” With its built-in speak tool and templated broadcast mechanism, the plugin supports features such as the model deciding what to say, status broadcasting, and a native macOS floating pet, and provides multiple TTS backends from cloud-based to local.

Core Features

The plugin implements voice interaction through the following mechanisms:

  1. Three speech modes

    • speak tool: The model autonomously decides what to say (e.g., results or progress), presented in natural language rather than reading out the full reply.
    • readReplies: Reads the assistant’s messages word by word (disabled by default), and is mutually exclusive with the speak mode.
    • Verbalizer: After a conversation turn ends, a lightweight LLM rewrites the reply into a single emotionally toned sentence (e.g., “Good news, all tests passed”), with streaming generation support.
  2. Status broadcasting

    • Provides templated announcements for events such as approvals, questions, turn start/end, and pending items.
    • Includes dedicated phrases for abnormal endings (interruption, blocking, errors, etc.).
  3. Session voice

    • Each session is automatically assigned a unique voice based on an FNV-1a hash, with a ±6% rate jitter added, removing the need for manual voice selection.
  4. Auditory assistance

    • Chimes: System-notification-style alert sounds (e.g., macOS glass sound), played before the voice.
    • Monitoring panel: Provides an SVG line chart view of token consumption, latency, and system statistics.
    • Native macOS Floating Pet: A native macOS floating window that displays the current utterance and pending items.
  5. TTS backends

    • Supports four backends: edge (default, cloud-based, Chinese and English), kokoro (local, English only), say (macOS offline), and fake (testing).

Installation and Activation

The installation process involves building and configuring links:

  1. Build the plugin
    Run the build command in the plugin directory:
    npm run build
This generates `lib/client.js` and the required configuration files.
  1. Build the native pet (optional, macOS only)
    If you need the native floating pet feature, you must build it with Swift:
    DEVELOPER_DIR=/Library/Developer/CommandLineTools swift build --package-path pet -c release
  1. Link it to the DSH configuration
    In DSH’s package.json, add the dependency and specify the entry:
    • Add a link: dependency in dependencies.
    • Specify entry in bundle to point to ./cordis.patch.yml.
    • Restart DSH Desktop.

Typical Usage

The plugin uses presets and configuration to provide voice feedback for different scenarios:

  1. Configure presets
    Configure three presets via the settings panel or API:

    • Immediate: speech rate +18%, announces every step.
    • Default: rate 0%, announces only when delivering results or encountering a block.
    • Minimal: rate -5%, volume -30, announces only approvals and questions, while the assistant remains silent (chimes only).
  2. Handling blocking events
    When the model cannot continue execution (e.g., when user input is required), a templated announcement prompts the user, avoiding silence while the model waits.

  3. Switching language
    The plugin supports full zh/en internationalization switching. UI text, broadcast phrases, and Verbalizer prompts are all controlled by the locale, and can be manually overridden in the settings panel instead of the auto-detected browser language.

  4. Verifying features
    Use the test script to verify TTS and routing:

    npm test

Use Cases and Considerations

This plugin is suitable for developers or users who need multitasking and rely on auditory monitoring to track task progress.

  • Network dependency: Using the edge backend requires an internet connection; kokoro uses local inference but supports English only.
  • Mode exclusivity: The readReplies and speak modes cannot be enabled at the same time.
  • System requirements: The native floating pet feature requires a macOS environment and successful compilation of the Swift code.
  • Permissions and security: The plugin runs with DSH process privileges. Before installation, please review the source code and license (MIT).

Conclusion

By separating model decisions from status templates, dsh-voice-mini provides DeepSeek Harness with a flexible auditory feedback layer. Whether using cloud TTS or a local offline solution, it extends agent interaction from the screen to speech, improving the user experience.

Project links:
* GitHub
* Community directory