Douyin Knowledge Collector
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md and install @user_ed25382d/douyin-knowledge-collector.
About this skill
Problem
Collecting industry knowledge from Douyin is often fragmented: videos contain practical opinions, but content is hard to search, transcribe, and organize. Manually downloading, converting audio, and cleaning Whisper output creates repetitive work, especially when Chinese transcription lacks punctuation, mixes traditional characters, or contains homophone errors. This skill provides a repeatable pipeline for engineers and analysts: input keywords, output transcripts and structured source material for a knowledge base.
How it works
The core flow is: Chrome CDP search → video URL extraction → audio or video download → Whisper transcription → quality reports and knowledge-base drafting. Key steps:
- Search and collection: connects to a logged-in Chrome session via Chrome DevTools Protocol, takes first-page results, uses random pauses, video intervals, and CDP retry behavior to reduce instability.
- Audio extraction: prefers DASH audio streams, falls back to full video when needed, and supports resumable downloads with retries for network interruptions.
- Transcription and cleanup: uses
openai-whisperfor Chinese transcription, guides punctuation output withinitial_prompt, and applies a built-in traditional-to-simplified mapping. - Quality grading: generates
quality_report.json, classifying transcripts ashigh,medium, orlowso longer and more usable outputs can be prioritized. - Knowledge drafting: merges texts under
transcripts/, removes duplicates, groups content by topic, and produces raw and refined knowledge-base files.
Boundaries
The skill is designed for Douyin only and does not support Bilibili, Xiaohongshu, or YouTube. It requires a logged-in Chrome session, Python, ffmpeg, and whisper. Whisper Chinese accuracy is about 85%, so final knowledge-base content still needs human review for homophone errors, domain terms, numbers, and uncertain items.
Use Cases
- Operations researchers collect Douyin videos by multiple keywords and transcribe them into comparable operational notes.
- Analysts batch-extract industry short-video audio and generate transcripts for competitor-positioning comparison.
- Researchers transcribe spoken Douyin videos, grade output quality, and merge content into a structured knowledge base.
- Strategy teams gather first-page Douyin opinion samples and output raw and refined knowledge-base files.
Best For
- Operations analysts who need to extract industry tactics from Douyin short videos weekly
- Competitive analysts who need to transcribe competitor video narration and organize key points
- Researchers who need to collect practical knowledge from Douyin topic videos
- Finance or strategy analysts who need cross-source opinion validation from Douyin videos
Related Skills
Fetches Baidu Hot Search Top 10 titles using web_fetch first, validates same-day data, and falls back to browser automation when stale.
Generates an evening A-share policy and trading opportunity daily report by collecting same-day index, policy, and capital data, then applying a fixed template to highlight beneficiary sectors, drivers, and price directions.
Maps natural-language TikTok requests to KeyAPI REST scenarios, validates endpoints against docs, and executes data queries and analysis.
Extract city-specified AI jobs from BOSS Zhipin, save CSV/table data, mark new postings, and summarize salary trends, application advice, and HTML reports.