AI Agent Hub
Back to skills
OpenClaw HearSpeak Voice Assistant icon

OpenClaw HearSpeak Voice Assistant

AI Agent Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md and install @user_aae27d93/openclawhearspeak.

About this skill

Problem addressed

In a local Linux terminal, voice assistants often depend on online wake-word detection, manual recording control, and shared session state that collides with a Web UI. openclawHearSpeak targets this workflow by chaining offline wake-word detection, automatic endpointing, independent conversation state, and spoken responses into one local voice pipeline.

How it works and limitations

  • Wake-up and capture: it uses sherpa-onnx for offline keyword detection, with “小右” as the default wake word. After startup, it listens in the background, starts recording on the wake word, and uses VAD to detect speech start and end. The docs describe ending capture about 1.5s after the user stops talking and timing out if no speech begins within 5s.
  • Recognition and dialogue: audio is sent to the SiliconFlow ASR service configured through SILICONFLOW_API_KEY, recognized as Mandarin Chinese text, and passed to the OpenClaw Agent. Responses are received via message subscription in a streaming style, using a clean independent session to avoid conflicts with the Web UI; periodic voice prompts are played while the gateway is processing.
  • Output and devices: final answers are synthesized with Edge-TTS for natural playback, and content such as emoji and icons is filtered out because it is unsuitable for speech. On first run, the assistant detects the microphone, calibrates background noise, downloads the KWS model, and generates prompt sounds.
    It is best suited to Python 3.8+, Linux, and ALSA environments with a USB microphone and speaker for Mandarin Chinese conversation. Keep the microphone and speaker apart to reduce echo, keep the room quiet during first run, start speaking promptly after wake-up, and avoid running multiple instances at once because that can cause device conflicts.

Use Cases

  • Set up an offline voice Q&A session on Linux, waking the assistant with “小右” and asking the Agent via a USB microphone.
  • Complete first-run microphone detection and noise calibration, then verify whether Edge-TTS can play back the answer.
  • Use an independent voice conversation session with an existing OpenClaw gateway so it does not pollute the Web UI session.
  • Tune silence-end, wake-up wait, and calibration duration during long runs to reduce early cutoffs or false wakes.

Best For

  • Linux engineers maintaining an OpenClaw gateway who want offline wake-word Chinese voice Q&A.
  • Technical authors building voice desktop prototypes who need microphone detection, noise calibration, and TTS playback checks.
  • Hardware test engineers validating meeting terminals who must check USB mics, speaker echo, and hot-plugging.
  • Backend engineers debugging Agent session isolation who want an independent voice session and streaming replies.