AI Agent Hub
Back to skills
Automatic Mixed-Cut Video Production Workflow icon

Automatic Mixed-Cut Video Production Workflow

Design & Media Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @user_51873b1f/jilaoshi-video-workflow into my AI assistant according to https://skillhub.cn/install/skillhub.md.

About this skill

What It Solves

When you already have a final Chinese script and voiced narration, mixed-cut video work often stalls on caption timing, clean 4:3 crops, source de-duplication, consistent cover identity, and draft sync. This workflow fixes the delivery target: a 2160x2880, 25 fps vertical MP4 with a black-gold 3:4 cover embedded as frame 0, one-line Chinese captions matched to the narration, plus a source manifest, shotlist, and QA reports.

How It Works

It treats the narration audio as the only timeline. ASR is used for word-level timestamps, then timing is realigned to the user's original script; displayed captions keep every spoken character and must match the script after punctuation and whitespace removal. It builds a duration-first shotlist, finds a unique source for each shot, and disallows looping, ping-pong, freeze extension, or self-concatenation. Clips are trimmed or cropped to clean 2160x1620 assets and assembled into a black 4:3 picture window. The cover is generated and validated, then used as frame 0. Final packaging uses ffmpeg / ffprobe for scaling, grading, mixing, and verification, targeting about -16 to -14 LUFS and true peaks no higher than -1 dBFS.

Boundaries

It assumes local ffmpeg, a transcription route that can produce timestamps, and image generation for the cover. Missing optional plugins should not block the final MP4. If the script or narration is missing and no safe fallback exists, the workflow should stop and request the asset. Sources should retain source pages, rights notes, and content hashes, not just filenames.

Use Cases

  • With a Chinese script and voiced audio ready, produce a vertical 4K MP4, cover frame 0, and QA reports.
  • Align narration ASR timestamps to the original script, generate non-overlapping one-line captions, and verify full-text exactness.
  • Plan shots with real-person video, prioritize named speeches or interviews, and avoid looping, duplicates, and filename-only checks.
  • Generate a black-gold 3:4 cover, inspect faces, text, and perspective, then embed it as the final video frame 0.

Best For

  • Video creators who have a Chinese script and narration and need editing, captions, cover, and delivery validation.
  • Content operators who need publishable vertical MP4 output with a source manifest and QA reports.
  • Editing engineers who want a repeatable `ffmpeg` and transcription workflow for `4:3` cropping and mixing.
  • Self-media editors who need real-person assets, license records, and a synced Jianying/CapCut draft.