Recapo Video Narration Editing
Paste the following prompt into your AI chat to install this skill:
Please install @user_e940e7fc/recapo-video-narration by following the guide at https://skillhub.cn/install/skillhub.md.
About this skill
Problem
Editing a film or series longer than two hours into a few-minute to half-hour narration usually stalls on scene selection, pacing, and audio-visual alignment. Watching the full source, drafting a script, generating TTS audio, and manually correcting the timeline is costly, and local dependencies can vary across machines.
How It Works
Recapo Video Narration Editing runs a fixed multi-agent state machine:
- Content understanding: clips are processed in 5-minute chunks to produce characters, a full-film story outline, and shot organization.
- Selection and scoring: NarrativeScoreAgent scores narrative value, SegmentFilterAgent filters segments, and EpisodeSelectAgent selects episodes with up to five timing corrections.
- Script and voice: SceneGrouperAgent groups scenes, EpisodeArrangeAgent preserves complete scenes and original audio silences, ScriptGenerationAgent writes scene-level narration, ScriptFixAgent refines word count for up to five rounds, then Doubao TTS generates audio.
- Assembly and validation: the skill adjusts selected ranges and subtitle timing using real TTS duration and original clip ratios, then completes synthesis and quality checks with FFmpeg.
It requires target durations to fall within 1-30 minutes and total output length to stay under the source length. Before running, it reads references/parameters.md, and also references/api-setup.md for credentials or API errors, then shows a parameter summary for confirmation. Local execution uses the user’s own Ali Bailian and Doubao keys, via credentials.py or environment variables; keys must not be written into task config, command lines, logs, or results. On failure, the working directory is retained, and the report includes the error category, completed stages, and recovery command instead of silently restarting.
Boundaries
This skill fits automatic narration editing for movies and series. It is not intended to bypass organization security controls or Windows application restrictions. Runtime for long videos depends on source length, model response, and local performance. If subtitle or narration timeline validation fails, the task must not be marked successful.
Use Cases
- Received a 2-hour film and need a 15-minute Chinese narration cut with scene silences and subtitles.
- Split a series into multiple 8-minute English shorts, generating narration scripts, TTS audio, and timing.
- Rework a long video with existing BGM into a styled narration version and validate audio-visual sync.
- After a local run fails, resume selection, TTS, and FFmpeg synthesis using the same task config.
Best For
- Video editors distributing film content who need long movies turned into multiple narration cuts.
- Content creators producing video accounts who need Chinese scripts and TTS output from full films.
- Engineers running local AI workflows who need credential setup and state-machine selection and synthesis.
- Post-production staff handling series episodes who need episode-level cuts, audio silences, and subtitle timing checks.
Related Skills
Excalidraw Wrap is an Excalidraw-focused wrapper, with tags for TypeScript, GitHub, and automation.
Calibrate vague brand inputs, expose contradictions, distill a brand core and positioning boundaries, then stress-test the result into an executable brand skeleton.
Generate localized Chinese brand names, naming directions, slogans, and risk checklists with reusable templates and trademark search reminders.
A beginner-friendly photo analysis tool that infers shooting parameters from visual features and suggests post-processing, optimization, and learning keywords.