Introduction¶
In the development of DeepSeek Harness (DSH), adding voice interaction capabilities to agents typically requires manually integrating TTS/ASR engines and building browser-side recording/playback UI. The dsh-voice plugin is designed specifically for interview preset scenarios and provides a turn-based voice channel. It abstracts away underlying engine details through Agent tools, enabling the interview flow to form a closed loop at the voice level.
Plugin Positioning¶
dsh-voice is a turn-based voice channel plugin that supports TTS/ASR engine abstraction. It provides voice interaction capabilities for interview scenarios, is maintained by its owner, and is licensed under MIT.
Core Features¶
The plugin provides voice capabilities through three Agent tools, with automatic browser-side recording and playback:
voice_speak: The Agent outputs voice, and the browser automatically plays it.voice_listen: Captures human voice input, transcribes it, and automatically sends it back to the Agent.voice_status: Engine health check; verifies TTS/ASR availability, queues, and recording status.
Engine Abstraction¶
The plugin allows independent engine selection for TTS and ASR. The auto mode probes available engines in the order qwen → mimo → local.
- qwen: Uses the native DashScope API and depends on the
DASHSCOPE_API_KEYenvironment variable. - mimo: Uses the MiMo-V2.5 API and depends on the
MIMO_API_KEYenvironment variable. ThebaseUrlcan be configured. - local: Supports local binaries and models such as sherpa-onnx, whisper.cpp, and piper.
The local ASR (sherpa-onnx) closed loop has been verified, supporting Chinese synthesis from macOS say and subsequent transcription. Audio is saved to ~/.dsh/voice by default.
Installation and Configuration¶
Run the installation command under the corresponding profile:
dsh plugin --profile web add @harness-flow/dsh-voice
After installation, restart the profile to apply the changes. The default configuration for the plugin line is neutral. It needs to be overridden in the profile’s cordis.patch.yml by the line id dsh-voice.
The example configuration is as follows:
- id: dsh-voice
config:
asr:
engine: auto
language: zh
tts:
engine: auto
voice: Cherry
audioDir: ~/.dsh/voice
listenTimeoutSec: 120
- Cloud engine credentials are read only from environment variables; configuration only declares the variable name.
- Local engines (such as sherpa-onnx) require the profile to install
sherpa-onnx-nodeand configure the model directory. - A local MiMo gateway can reuse the same implementation by configuring
baseUrlto point to your local address.
Interviewer Preset Usage¶
The plugin handles the voice path (host-side services + browser UI), while the preset is responsible only for the “interviewer” role and workflow.
It is recommended to define the interviewer Persona in the preset, requiring it to ask only one question at a time, read the question using voice_speak, and immediately call voice_listen to wait for the answer. The following is an example configuration fragment:
- id: persona
name: '@deepseek-ai/dsh-persona'
config:
text: >-
你是一位专业的面试官,正在通过语音进行一场结构化面试。
一次只提一个问题:用 voice_speak 朗读问题,随后立刻调用 voice_listen
等待候选人作答;根据 transcript 追问或进入下一题。
流程:开场寒暄 → 自我介绍 → 技术/项目深挖 → 行为面 →
候选人提问 → 结束语与后续安排。全程专业、中立、鼓励,不做主观臆断;
面试结束后输出评分表(各维度 1–5 分 + 一句话依据)与录用建议。
Version Notes and Limitations¶
The current version is v0.1, with the following design trade-offs:
- Turn-based priority:
voice_speakblocks until synthesis completes to ensure ordering; full-duplex barge-in will be supported in later versions. - Single-client bridging: Only single-client bridging is supported; listen sessions are managed in memory, and the queue is automatically discarded after the page is closed.
- No session persistence:
voice/*session events are not persisted; only chat records retain text transcriptions. - Localized failures: Engine probe failures affect only that engine; if browser playback is blocked, the system waits for a user gesture to retry.
Summary¶
dsh-voice addresses the engine integration and interaction logic challenges for building voice interview scenarios in DeepSeek Harness. It provides clear Agent tool interfaces and configurable TTS/ASR engine abstraction, making it suitable for developers who need to build voice-based interview workflows.