OpenClaw HearSpeak Voice Assistant
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md and install @user_aae27d93/openclawhearspeak.
About this skill
Problem addressed
In a local Linux terminal, voice assistants often depend on online wake-word detection, manual recording control, and shared session state that collides with a Web UI. openclawHearSpeak targets this workflow by chaining offline wake-word detection, automatic endpointing, independent conversation state, and spoken responses into one local voice pipeline.
How it works and limitations
- Wake-up and capture: it uses
sherpa-onnxfor offline keyword detection, with “小右” as the default wake word. After startup, it listens in the background, starts recording on the wake word, and uses VAD to detect speech start and end. The docs describe ending capture about1.5safter the user stops talking and timing out if no speech begins within5s. - Recognition and dialogue: audio is sent to the SiliconFlow ASR service configured through
SILICONFLOW_API_KEY, recognized as Mandarin Chinese text, and passed to theOpenClaw Agent. Responses are received via message subscription in a streaming style, using a clean independent session to avoid conflicts with the Web UI; periodic voice prompts are played while the gateway is processing. - Output and devices: final answers are synthesized with
Edge-TTSfor natural playback, and content such asemojiand icons is filtered out because it is unsuitable for speech. On first run, the assistant detects the microphone, calibrates background noise, downloads the KWS model, and generates prompt sounds.
It is best suited toPython 3.8+,Linux, andALSAenvironments with a USB microphone and speaker for Mandarin Chinese conversation. Keep the microphone and speaker apart to reduce echo, keep the room quiet during first run, start speaking promptly after wake-up, and avoid running multiple instances at once because that can cause device conflicts.
Use Cases
- Set up an offline voice Q&A session on Linux, waking the assistant with “小右” and asking the Agent via a USB microphone.
- Complete first-run microphone detection and noise calibration, then verify whether Edge-TTS can play back the answer.
- Use an independent voice conversation session with an existing OpenClaw gateway so it does not pollute the Web UI session.
- Tune silence-end, wake-up wait, and calibration duration during long runs to reduce early cutoffs or false wakes.
Best For
- Linux engineers maintaining an OpenClaw gateway who want offline wake-word Chinese voice Q&A.
- Technical authors building voice desktop prototypes who need microphone detection, noise calibration, and TTS playback checks.
- Hardware test engineers validating meeting terminals who must check USB mics, speaker echo, and hot-plugging.
- Backend engineers debugging Agent session isolation who want an independent voice session and streaming replies.
Related Skills
Provides Claw with character-library selection, switching, saving, and global SOUL.md style sync for role-based conversation.
An AIONE Agentic AI Infrastructure SDK wrapper for building production AI agents with memory, skills, workflows, and hooks.
A Python/TypeScript SDK wrapper for the DeepSeek-Reasonix native AI coding agent, with prefix-cache support.
A browser automation tool for analysts, operators, and developers that locates elements, fills forms, extracts structured content, and supports no-code scheduling and export.