AI Agent Hub
Back to skills
Douyin Video Transcript Extractor icon

Douyin Video Transcript Extractor

Knowledge Management Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @user_75818cbb/douyin-video-word according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

Douyin videos often contain spoken content that is hard to capture as plain text. This skill turns v.douyin.com / douyin.com video links into editable transcripts for archiving, search, and further processing.

How It Works

The skill is built around speech transcription: it uses yt-dlp to obtain audio, ffmpeg for format handling, and faster-whisper to convert speech to text. The first run checks dependencies and caches paths, while later runs start faster. It supports --lang auto/zh/yue/en/ja/ko, --model tiny/base/small/medium, --timestamps, --output, and --cookies. It corrects common Whisper homophone errors using context and includes some built-in correction for Buddhist terminology. When a request returns 403, it falls back through audio streams, format lists, and Playback API formats.

Boundaries

Most public videos can be processed without extra setup; restricted or login-gated content may require a cookie file. Higher-accuracy models use more time and resources. For Cantonese, dialects, noisy audio, or technical terms, use a larger model and review the output manually. Transcripts may still contain subtle speech-to-text errors, so they should not be treated as exact quotations.

Use Cases

  • Turn a shared Douyin talk video into an editable text draft for knowledge base archiving.
  • Choose language and Whisper model, then output transcript with timestamps when organizing video insights.
  • When public video audio fails with 403, let the script retry with fallback download strategies.
  • Run extraction with a local cookie file when login-gated or restricted Douyin content is needed.

Best For

  • Content operators who turn Douyin spoken videos into reusable text assets
  • Research assistants compiling multilingual video insights and checking terminology
  • Knowledge managers processing restricted videos and producing timestamped transcripts
  • Independent developers who need searchable text from spoken content