AI Agent Hub
Back to skills
FunASR Audio Transcription icon

FunASR Audio Transcription

Design & Media Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please install @user_e8cce942/funasr-transcriber according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

When the source material is meeting recordings, interview MP3s, or video soundtracks and you need a punctuated transcript, FunASR Audio Transcription converts speech in audio or video files into text. It targets the case where a voice file exists but searchable text is missing, not real-time captioning or video content understanding.

How it works

The skill provides a configurable transcription workflow based on FunASR. It can process WAV, MP3, M4A, and FLAC audio files; video files require audio extraction with FFmpeg. By default, --model=sensevoice selects SenseVoiceSmall with about 234M parameters, suitable for Chinese, Cantonese, English, Japanese, and Korean, with emotion and event detection. For broader language coverage and higher accuracy, --model=nano selects Fun-ASR-Nano-2512 with about 800M parameters, covering 31 languages and dialects. VAD and punctuation sub-models are loaded automatically when needed: fsmn-vad performs voice activity detection, and ct-punc restores Chinese and English punctuation.

Boundaries

The skill focuses on speech-to-text and does not perform visual description, chapter summarization, or translation output. Choose the model based on language coverage, parameter size, and accuracy needs. If the audio has complex background noise or the language is outside the supported list, confirm model support before deciding between SenseVoice and Nano.

Use Cases

  • Convert interview MP3 files into punctuated Chinese transcripts for searchable meeting notes.
  • Extract video lecture audio with FFmpeg, then generate English or Chinese text transcripts.
  • Process multiple meeting WAV files with VAD and punctuation restoration to create readable transcripts.
  • Transcribe dialect-heavy or less common language audio with the Nano model for higher accuracy.

Best For

  • Podcast editors who turn weekly interview MP3s into searchable written transcripts.
  • Course producers who convert video course audio into text drafts for later editing.
  • Audio engineers processing multiple meeting WAV files with VAD and punctuation restoration.
  • Localization researchers handling dialects or less common languages and seeking higher accuracy.