AI Agent Hub
Back to skills
Enhanced Video Downloader & Transcriber icon

Enhanced Video Downloader & Transcriber

Data Analysis Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_430ba93f/video-downloader-enhanced.

About this skill

What It Solves

Short-video platform assets are often scattered across web pages, publish copy, audio, and metadata. When building content research, script analysis, or ASR corpora, manually saving the video, copying titles and descriptions, extracting audio, and generating transcripts can quickly become inconsistent and hard to audit.

How It Works

The skill turns a supported platform URL into a local source-material folder with a stable artifact contract:

  • the downloaded video file;
  • post_caption.txt with title, description, and hashtags;
  • audio.wav / audio.m4a / audio.mp3 when audio extraction is used;
  • transcript.txt, plus transcript.srt for whisper.cpp-based workflows;
  • metadata.json with platform, author, duration, resolution, and download or ASR status.

It recognizes Douyin, Bilibili, YouTube, and Xiaohongshu links, selects the appropriate provider route, and can transcribe audio with whisper.cpp, openai-whisper, SiliconFlow, or none. Options such as --asr-prompt can bias transcription toward domain terms or request Simplified Chinese output. Output folder names include date, a trimmed title summary, platform, and author, with deduplication suffixes when a folder already exists.

Boundaries

WeChat Channels is only a reserved extension point, so direct download should not be claimed until a provider module is implemented and tested. For those links, say plainly that it is unsupported and let the user obtain the local file through an external tool first, then use the skill for local ASR or downstream processing. Use it only for material the user owns, is permitted to download, or can lawfully archive, and do not bypass DRM, paid access controls, or platform restrictions.

Use Cases

  • Archive Douyin, Bilibili, YouTube, and Xiaohongshu links into local folders with video, captions, and metadata.json.
  • Extract audio from short videos and generate transcript.txt plus timestamped transcript.srt with whisper.cpp.
  • Batch-collect video titles, descriptions, authors, duration, and resolution for script analysis or competitor review.
  • Check whisper.cpp, openai-whisper, and SiliconFlow availability under --asr auto before running speech transcription.

Best For

  • Content researchers who need to convert multi-platform short videos into local corpora with titles, descriptions, and metadata.
  • ASR engineers who need batch Chinese transcripts and SRT output, with local or cloud backend switching.
  • Short-video operators who need to archive competitor videos, author info, and publish copy for topic analysis.
  • Data engineers who need standardized, auditable folders containing video, audio, transcript text, and JSON metadata.