Introduction

DeepSeek Harness (DSH) provides plugin-based extensibility. When developing on the Web, adding voice interaction to the chat interface usually requires writing a backend service or integrating an external voice SDK. This plugin uses the browser-native Web Speech API to implement purely client-side speech recognition and text-to-speech (TTS) functionality, without server-side support or API Key configuration.

Plugin Overview

  • Name: lak321/dsh-voice-chat
  • One-line description: DeepSeek Harness voice chat plugin that supports voice input and automatic/manual read-aloud replies.
  • Maintainer: lak321
  • Category: Client
  • License: MIT

Core Features

  • Voice input: Performs speech recognition using the Web Speech API, supporting Mandarin, Cantonese, Taiwanese Mandarin, and English. Recognition results are displayed in real time in the input box and can be edited before sending.
  • Reply read-aloud: Integrates the browser TTS engine to read AI replies aloud. Provides automatic read-aloud (enabled by default), using message sequence numbers to distinguish new and old messages and avoid rereading historical messages.
  • Voice settings: Allows selecting a voice, adjusting rate (0.5~2.0), and pitch (0~2.0).
  • Text cleaning: Automatically removes Markdown markers (such as bold, italics, code blocks, headings, etc.) to prevent TTS from reading noise such as “asterisk asterisk”.
  • Data persistence: All settings are stored via localStorage and remain after page refresh.
  • Architectural features: Pure client-side plugin with no backend dependency and no API Key requirement.

Installation and Activation

This plugin is compatible only with DSH Profile: web.

Install using the official command (git source, artifacts are already in the repository, no build required):

dsh plugin --profile web add "github:lak321/dsh-voice-chat#<commit>&path:/"

After installation, restart the DSH process (for example, run npx -y @deepseek-ai/dsh web).

Note: The old manual method of copying the client/ directory and injecting cordis.patch.yml is deprecated.

Typical Usage

  1. Voice input: Click the button next to the input box and speak. The recognized text is automatically filled into the input box and can be edited and sent directly.
  2. Reply read-aloud: A read-aloud button appears next to each AI reply. Click it to read that message aloud. Automatic read-aloud is enabled by default, and new replies are read automatically when generated.
  3. Settings adjustment: In the settings panel, you can disable automatic read-aloud, select a voice, or adjust the rate and pitch.

Development and Caveats

  • React Hook conventions: If developing with React, props.useSession must be called at the top level of the component function body. It cannot be called inside event handlers such as onClick, otherwise an error will occur.
  • Export requirement: Client plugins must export exports.inject to use ctx.slots.
  • Runtime mechanism: The plugin registers slots to conversation.input.left and conversation.chat.assistant-actions via ctx.slots.inject, and implements the core logic through SpeechRecognition and speechSynthesis.

Conclusion

This plugin provides lightweight voice interaction capabilities for the DSH Web interface and is suitable for scenarios that require direct speech functionality in the browser.