Bilibili Video Analysis
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @user_630f04d9/bilibili-browse-videos.
About this skill
Problem
Bilibili videos usually separate the visual track, audio track, and subtitles. Manually clipping clips, transcribing speech, and summarizing frames is fragmented. This skill targets short-video analysis and turns a link into a readable, reviewable summary.
How It Works
- Input: accepts short links, full Bilibili URLs, or raw
BVIDvalues. - Download and preprocessing: resolves download URLs with
bilibili-api-python, downloads m4s streams withaiohttp, converts audio to16kHz mono WAV, and extracts keyframes withffmpeg. - Subtitles: runs local
openai-whisperwith themediummodel to produce Chinese time-stamped subtitles. - Frame analysis: selects 6-11 frames based on duration and analyzes them with the image tool.
- Summary: combines subtitles, frames, and visual notes into a structured summary, then sends it to Feishu via the message tool.
Boundaries and Notes
The skill depends on ffmpeg, a cached Whisper model, and local storage, and is aimed at short videos; a 60-second clip takes roughly 3-4 minutes. Downloads require a Referer, and file names use BVID to avoid special characters. Whisper may output Traditional Chinese; conversion is optional. It does not make copyright determinations and does not guarantee stable access to every Bilibili video.
Use Cases
- Summarize a 60-second Bilibili product explainer with Chinese subtitles and visual key points.
- Resolve a short link to a BVID, download audio and video, generate subtitles, and extract keyframes.
- Combine subtitles, keyframes, and visual notes into a structured summary sent to Feishu.
- Check whether spoken points and on-screen content in a short clip align, producing a reviewable digest.
Best For
- Content operators who turn Bilibili short clips into readable summaries
- Product managers who verify that spoken points and visual content align
- Analysts extracting key points from competitor videos for archiving
- Applied AI engineers needing local subtitles and keyframe analysis
Related Skills
Convert Chinese or natural-language requests into paste-ready English image prompts for ChatGPT's web UI, covering generation, editing, multi-image references, and exact text without performing image generation.
A creative AI image workflow for style transfer, reference-based creation, scene replacement, series expansion, material conversion, era shifts, composition rework, lineart conversion, and cartoonization.
Translates star-inspired football memories and fan resonance into compliant original IP poster concepts while avoiding real names, likenesses, official badges, and event logos.
Extracts real web page colors, typography, components, layouts, and interaction styles with a Playwright script to produce a verifiable design guideline.