DeepSeek Harness (DSH) Web interface does not have voice output by default. make-dsh-voice is a lightweight plugin designed for this scenario. It calls the Alibaba Cloud Bailian CosyVoice model to give the agent a voice and provides a persistent player and configuration entry in the interface. This plugin is maintained by the developer blueperformer and aims to solve the lack of voice interaction in the Web UI.

Core Features

This plugin mainly provides the following capabilities:

  • voice_speak tool: The agent can invoke this tool to synthesize speech using a registered voice cloning model.
  • Voice bar: A single-line player is displayed below the editor for playing the latest synthesized audio clip.
  • Voice settings page: A new “Voice” tab is added to the settings menu, supporting configuration of the API Key, model, voice ID, response strategy (whether to read every sentence), output length, and output directory.
  • Startup sound: After the page is loaded, a specified sound is played when the user performs the first interaction (click, input, etc.). This feature can be toggled on or off.

Installation and Enablement

The installation process consists of three steps and is completed via CLI commands.

  1. Install the plugin
    Use the dsh plugin command to add the repository. Make sure to use the web profile.
    dsh plugin --profile web add github:blueperformer/make-dsh-voice
  1. Restart Harness
    The plugin is mounted through the dsh.bundle mechanism, so the DSH process must be restarted to activate it.

  2. Configure the API Key
    Open DSH settings, find the “Voice” page, and enter the Alibaba Cloud Bailian API Key. After configuration, the key is retained locally only; the browser side will not receive the key and will only show a “Configured” status.

Typical Usage

After installation, you can use it in the following ways:

  • Set the startup sound
    Copy a prepared MP3 audio file to the plugin’s assets directory and rename it to boot.mp3. After the page is loaded, the system will attempt to play the audio when it detects the user’s first interaction.

  • Standalone voice synthesis
    The plugin includes a Python script synthesize.py, which can be used for local testing or standalone invocation.

    python synthesize.py --text "hello"

Environment Requirements and Notes

The plugin requires a backend Python service, so note the following environment requirements:

  • Operating system: Currently only tested in a Windows 11 environment; Linux and macOS have not been verified.
  • DeepSeek Harness: The Web profile of DeepSeek Harness (dsh web) must be used.
  • Dependencies:
    • pnpm must be installed and available in PATH.
    • Python 3 and the Alibaba Cloud DashScope SDK must be installed (pip install dashscope).
    • ffmpeg is optional and is used to process silence before and after the audio. If it is not installed, this step will be skipped.

Uninstall

To remove the plugin, run the following command:

dsh plugin --profile web remove dsh-voice

After uninstalling, the plugin will be removed from the configuration, but any voice configuration section retained in the settings file is usually not cleaned up automatically.