MiMo V2.5 TTS
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md and install @user_a5827f21/mimo-v2-5-tts.
About this skill
Problem
Calling a TTS API often only solves “read this text.” When the same line needs different tones, dialects, singing, or more natural pauses, teams may end up rewriting prompts, wiring APIs, and tuning parameters by hand. mimo-v2.5-tts targets Chinese and English speech generation by organizing model selection, style control, and audio output into a script-based workflow.
How It Works
It wraps three MiMo V2.5 entry points:
- mimo-v2.5-tts: preset voices, Chinese/English, and singing
- mimo-v2.5-tts-voicedesign: generates a new voice from a short description
- mimo-v2.5-tts-voiceclone: clones a voice from an mp3/wav sample
Natural-language control is passed via --context, describing character, scene, and delivery cues in a “director mode.” Inline style or audio tags can also be embedded in the text, such as (singing), (Northeastern Chinese), or [pause], to adjust tone, emotion, and breathing inside a sentence. Long text is usually generated in one pass; only when it exceeds about 2500 characters should you split by paragraph and join the audio with ffmpeg. For Feishu delivery, the skill can run the audio-send script to transcode, upload, and deliver an audio message.
Limits
Preset voices require an explicit voice ID; voice-clone samples must be Base64-encoded under 10 MB and use mp3/wav; and TTS output is stochastic, so important lines should be generated multiple times and selected.
Use Cases
- Prepare a Chinese podcast intro and generate a pre-picked voice with pauses and soft laughter using audio tags.
- Break product copy into multiple narrated roles and generate distinct host-like voices using voice design.
- Synthesize an English novel chapter into an audio segment and control pacing, breath, and emphasis in director mode.
- Send the generated WAV to a Feishu test group as an audio message and verify it is not delivered as a normal file.
Best For
- Audio editors producing podcasts who need to generate multi-emotion Chinese narration with consistent tone.
- Frontend engineers building product demos who want copy converted into speech with pauses and emphasis.
- Independent audio creators who want to design a unique voice from a text description and reuse it later.
- Automation engineers delivering Feishu messages who need TTS audio sent as an audio message, not a file attachment.
Related Skills
Convert Chinese or natural-language requests into paste-ready English image prompts for ChatGPT's web UI, covering generation, editing, multi-image references, and exact text without performing image generation.
A creative AI image workflow for style transfer, reference-based creation, scene replacement, series expansion, material conversion, era shifts, composition rework, lineart conversion, and cartoonization.
Translates star-inspired football memories and fan resonance into compliant original IP poster concepts while avoiding real names, likenesses, official badges, and event logos.
Extracts real web page colors, typography, components, layouts, and interaction styles with a Playwright script to produce a verifiable design guideline.