AI Agent Hub
Back to skills
Local VibeVoice Text-to-Speech icon

Local VibeVoice Text-to-Speech

Knowledge Management Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Install @org-zzl567pv/cs-obsidian according to https://skillhub.cn/install/skillhub.md.

About this skill

The problem it solves

Local text-to-speech setups often get stuck on model downloads, dependency versions, device compatibility, and multi-speaker script formatting. This skill wraps the local VibeVoice-1.5B inference path into an executable script, so the goal is not a cloud API call but generating .wav audio inside an existing uv environment.

How it works

The skill directory includes setup.sh, tts_generate.py, a virtual environment, nine bundled voices, and an output directory. First use creates a Python 3.11 virtual environment and installs vibevoice from GitHub; later runs use the locally cached model. Generation starts by classifying the input: plain text is treated as a single speaker, while multi-speaker scripts must use the Speaker N: ... format, with N from 1 to 4. It then selects voice presets such as the default zh-Xinran_woman, Chinese male zh-Bowen_man, or English en-Alice_woman. Finally, the skill's own virtual environment executes the script and writes output to outputs/tts_TIMESTAMP.wav.

Supported arguments include --text, --txt_path, --speaker_names, --output, --device, --seed, --ddpm_steps, --voices_dir, and --checkpoint_path. On CUDA, it uses bfloat16 and flash_attention_2, falling back to sdpa when unavailable. On Mac MPS, it uses float32 and sdpa, with roughly 6 GB of unified memory needed. The first model download is about 3 GB, and the model supports up to about 90 minutes of audio.

Boundaries and cautions

  • Short text can use --text; multi-line or long passages should be written to a .txt file and passed via --txt_path to avoid shell newline issues.
  • Chinese text works better with English punctuation such as , and ., reducing pronunciation instability.
  • Multi-speaker scripts must follow the strict format and support at most four speakers.
  • Generation speed is roughly 2 to 5 times real time on Apple Silicon, so 30 seconds of audio may take 60 to 150 seconds.
  • Custom voices are supplied through --voices_dir with .wav files; fine-tuned adapters can be loaded through --checkpoint_path.

Use Cases

  • Generate four-speaker podcast WAV files locally instead of sending sensitive scripts to cloud TTS.
  • Debug the 1.5B model on Apple Silicon, run MPS single-speaker speech, and check generation time.
  • Convert teaching scripts with Speaker 1 and Speaker 2 tags into classroom audio at fixed paths.
  • Generate character dialogue with custom .wav voices and LoRA checkpoints, then compare seed settings.

Best For

  • Audio engineers producing podcasts who need local multi-speaker TTS and fixed WAV output.
  • Researchers fine-tuning VibeVoice LoRA who need checkpoint loading and reproducible seeds.
  • Full-stack engineers prototyping voice on macOS MPS or CUDA and testing voice parameters.
  • Content teams handling sensitive scripts who need offline Chinese dialogue with output paths.