Local VibeVoice Text-to-Speech
Paste the following prompt into your AI chat to install this skill:
Install @org-zzl567pv/cs-obsidian according to https://skillhub.cn/install/skillhub.md.
About this skill
The problem it solves
Local text-to-speech setups often get stuck on model downloads, dependency versions, device compatibility, and multi-speaker script formatting. This skill wraps the local VibeVoice-1.5B inference path into an executable script, so the goal is not a cloud API call but generating .wav audio inside an existing uv environment.
How it works
The skill directory includes setup.sh, tts_generate.py, a virtual environment, nine bundled voices, and an output directory. First use creates a Python 3.11 virtual environment and installs vibevoice from GitHub; later runs use the locally cached model. Generation starts by classifying the input: plain text is treated as a single speaker, while multi-speaker scripts must use the Speaker N: ... format, with N from 1 to 4. It then selects voice presets such as the default zh-Xinran_woman, Chinese male zh-Bowen_man, or English en-Alice_woman. Finally, the skill's own virtual environment executes the script and writes output to outputs/tts_TIMESTAMP.wav.
Supported arguments include --text, --txt_path, --speaker_names, --output, --device, --seed, --ddpm_steps, --voices_dir, and --checkpoint_path. On CUDA, it uses bfloat16 and flash_attention_2, falling back to sdpa when unavailable. On Mac MPS, it uses float32 and sdpa, with roughly 6 GB of unified memory needed. The first model download is about 3 GB, and the model supports up to about 90 minutes of audio.
Boundaries and cautions
- Short text can use
--text; multi-line or long passages should be written to a.txtfile and passed via--txt_pathto avoid shell newline issues. - Chinese text works better with English punctuation such as
,and., reducing pronunciation instability. - Multi-speaker scripts must follow the strict format and support at most four speakers.
- Generation speed is roughly 2 to 5 times real time on Apple Silicon, so 30 seconds of audio may take 60 to 150 seconds.
- Custom voices are supplied through
--voices_dirwith.wavfiles; fine-tuned adapters can be loaded through--checkpoint_path.
Use Cases
- Generate four-speaker podcast WAV files locally instead of sending sensitive scripts to cloud TTS.
- Debug the 1.5B model on Apple Silicon, run MPS single-speaker speech, and check generation time.
- Convert teaching scripts with Speaker 1 and Speaker 2 tags into classroom audio at fixed paths.
- Generate character dialogue with custom .wav voices and LoRA checkpoints, then compare seed settings.
Best For
- Audio engineers producing podcasts who need local multi-speaker TTS and fixed WAV output.
- Researchers fine-tuning VibeVoice LoRA who need checkpoint loading and reproducible seeds.
- Full-stack engineers prototyping voice on macOS MPS or CUDA and testing voice parameters.
- Content teams handling sensitive scripts who need offline Chinese dialogue with output paths.
Related Skills
Turns Taleb's antifragility, barbell strategy, and via negativa into executable prompts for risk analysis in career, investing, health, and other domains.
Keyword news search over ZAKER’s corpus with optional time filtering, returning up to 20 titles, authors, and summaries.
A generic knowledge review skill that generates role-based checklists, rrule-based reminders, and on-demand reviews without storing personal data.
A knowledge-management skill for layered retrieval across information sources.