FunASR Audio Transcription
Paste the following prompt into your AI chat to install this skill:
Please install @user_e8cce942/funasr-transcriber according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem
When the source material is meeting recordings, interview MP3s, or video soundtracks and you need a punctuated transcript, FunASR Audio Transcription converts speech in audio or video files into text. It targets the case where a voice file exists but searchable text is missing, not real-time captioning or video content understanding.
How it works
The skill provides a configurable transcription workflow based on FunASR. It can process WAV, MP3, M4A, and FLAC audio files; video files require audio extraction with FFmpeg. By default, --model=sensevoice selects SenseVoiceSmall with about 234M parameters, suitable for Chinese, Cantonese, English, Japanese, and Korean, with emotion and event detection. For broader language coverage and higher accuracy, --model=nano selects Fun-ASR-Nano-2512 with about 800M parameters, covering 31 languages and dialects. VAD and punctuation sub-models are loaded automatically when needed: fsmn-vad performs voice activity detection, and ct-punc restores Chinese and English punctuation.
Boundaries
The skill focuses on speech-to-text and does not perform visual description, chapter summarization, or translation output. Choose the model based on language coverage, parameter size, and accuracy needs. If the audio has complex background noise or the language is outside the supported list, confirm model support before deciding between SenseVoice and Nano.
Use Cases
- Convert interview MP3 files into punctuated Chinese transcripts for searchable meeting notes.
- Extract video lecture audio with FFmpeg, then generate English or Chinese text transcripts.
- Process multiple meeting WAV files with VAD and punctuation restoration to create readable transcripts.
- Transcribe dialect-heavy or less common language audio with the Nano model for higher accuracy.
Best For
- Podcast editors who turn weekly interview MP3s into searchable written transcripts.
- Course producers who convert video course audio into text drafts for later editing.
- Audio engineers processing multiple meeting WAV files with VAD and punctuation restoration.
- Localization researchers handling dialects or less common languages and seeking higher accuracy.
Related Skills
Convert Chinese or natural-language requests into paste-ready English image prompts for ChatGPT's web UI, covering generation, editing, multi-image references, and exact text without performing image generation.
A creative AI image workflow for style transfer, reference-based creation, scene replacement, series expansion, material conversion, era shifts, composition rework, lineart conversion, and cartoonization.
Translates star-inspired football memories and fan resonance into compliant original IP poster concepts while avoiding real names, likenesses, official badges, and event logos.
Extracts real web page colors, typography, components, layouts, and interaction styles with a Playwright script to produce a verifiable design guideline.