Foreword

DeepSeek Harness (DSH) adopts the “everything is a plugin” philosophy, allowing users to extend functionality in the Web interface. For developers who are accustomed to voice interaction or need to keep their hands free while handling multiple tasks, typing is often not the most efficient input method. The dsh-voice-talk plugin solves this problem by providing a full-screen voice call layer that lets users talk directly to the AI through a microphone, creating a closed-loop experience of speaking and listening while responses are generated and read aloud.

Core Capabilities

The plugin mainly provides the following capabilities:
1. Full-screen call mode: Click the microphone icon to enter the full-screen call layer. The interface supports bilingual Chinese and English and automatically adapts to light and dark themes.
2. Streaming interaction: Supports voice input; the AI response is read aloud while it is being generated.
3. Voice adjustment: Supports speech-speed adjustment and voice switching. The default system voice is used.
4. Security: The API key is stored on the local machine only and is never uploaded. The browser cannot directly access the key, and voice requests are forwarded by the host proxy.
5. Open source: Released under the MIT License.

Installation and Prerequisites

Before installing, make sure the host environment meets the requirements.

Prerequisites
* DeepSeek Harness Web Profile version: 0.1.0-rc.7+ (compatible with both the 0.1.0-rc.x and 0.1.5-rc.x host series).
* Browser: Chrome or Edge.
* Network: Internet access is required.

Installation steps
1. Open a terminal and run the installation command:

    dsh plugin --profile web add dsh-voice-talk
  1. After installation, restart the Web process.
    dsh web

Configuration and Usage

After installing the plugin, configure it before using it.

1. Configure voice recognition
Voice recognition requires a cloud-based engine. Users need to apply for an API Key on Alibaba Cloud Bailian and configure it in DSH.
* Log in to the Alibaba Cloud Bailian console and apply for an API Key.
* In the DSH interface, go to Settings → Plugins → Voice Talk.
* Enter the API Key in the Qwen Settings. The model and endpoint are built in and become available after recharge.

2. Configure voice synthesis
* Default mode: Uses the system voice and is free.
* Natural voice: For a more natural voice, switch to the Qwen voice. It is pay-as-you-go (approximately CNY 1 per 10,000 characters).

3. Usage
1. Find the microphone icon next to the input box on the chat page or workbench.
2. Click the icon to enter the full-screen call layer, then speak directly.
3. AI responses are automatically generated and read aloud.
4. You can mute or minimize the call layer during a call.
5. To exit the call layer, click the red hang-up button or press Esc.

Settings

Settings on the settings page are persistent defaults; settings that apply only to the current call should be adjusted in the call layer.

Setting Description Default
Interrupt playback when speaking Whether speaking interrupts the ongoing playback (headphones recommended) Off
Auto-send pause duration (seconds) How long to pause after speaking before automatically submitting to the AI 2
Voice waveform effect Call-layer waveform style: Rhythm / Ripple Ripple
Speech speed Playback speed 1.0x

Notes

  1. Browser compatibility: Chrome and Edge are currently supported.
  2. Permissions: Microphone permission must be granted before use.
  3. Key security: The plugin follows the principle of least privilege. Keys are stored locally only, requests are forwarded by the host proxy, and browser-side code cannot access the keys.
  4. Version dependency: Use Web Profile version 0.1.0-rc.7 or later.
  5. Feedback channels: If you encounter issues, submit an issue on GitHub or leave a comment under the plugin marketplace card.

Summary

dsh-voice-talk uses a full-screen voice call mode to transform the DSH Web interface from a simple typing tool into a multimodal interaction terminal. It combines the free convenience of system voices with the natural experience of Qwen voices, making it suitable for development scenarios that require hands-free, rapid voice interaction.