AI Agent Hub
Back to skills
Video to Markdown Transcriber icon

Video to Markdown Transcriber

Design & Media Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Install @user_fe8721be/video-to-md according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

Video material often lives in URLs, local files, and platform-specific pages, while engineers need searchable text for review, indexing, or downstream processing. This skill focuses on turning supported sources into transcripts: bilibili.com, YouTube, TikTok/Douyin, Twitter/X via yt-dlp, and local files such as mp4, wav, m4a, webm, mkv.

How It Works

It uses yt-dlp to acquire media, ffmpeg to process audio, and faster-whisper for speech recognition. Key controls include:
- url: video URL or local file path
- -m, --model: choose tiny, base, small, medium, or large
- -l, --language: force a language code, otherwise auto-detect
- -o, --output: write a file or print to terminal
- --keep-files: retain downloaded media files

Bilibili inputs require auth values such as SESSDATA, bili_jct, and buvid3, which can be supplied through config or command-line arguments. These values are login credentials and should not be shared. tiny and base are faster and suitable for quick drafts; medium and large are slower but more accurate for final transcripts.

Use Cases

  • Export Bilibili lecture video transcripts for meeting notes.
  • Convert local mp4 interview audio into searchable text.
  • Generate YouTube tech-share transcripts using medium model.
  • Transcribe TikTok videos as English text for asset cleanup.

Best For

  • Knowledge ops staff who turn course videos into editable notes
  • Technical editors extracting key points from shared talks
  • Content engineers converting short video speech into searchable text
  • ML engineers validating Whisper behavior on local video files