Automatic Mixed-Cut Video Production Workflow
Paste the following prompt into your AI chat to install this skill:
Please install @user_51873b1f/jilaoshi-video-workflow into my AI assistant according to https://skillhub.cn/install/skillhub.md.
About this skill
What It Solves
When you already have a final Chinese script and voiced narration, mixed-cut video work often stalls on caption timing, clean 4:3 crops, source de-duplication, consistent cover identity, and draft sync. This workflow fixes the delivery target: a 2160x2880, 25 fps vertical MP4 with a black-gold 3:4 cover embedded as frame 0, one-line Chinese captions matched to the narration, plus a source manifest, shotlist, and QA reports.
How It Works
It treats the narration audio as the only timeline. ASR is used for word-level timestamps, then timing is realigned to the user's original script; displayed captions keep every spoken character and must match the script after punctuation and whitespace removal. It builds a duration-first shotlist, finds a unique source for each shot, and disallows looping, ping-pong, freeze extension, or self-concatenation. Clips are trimmed or cropped to clean 2160x1620 assets and assembled into a black 4:3 picture window. The cover is generated and validated, then used as frame 0. Final packaging uses ffmpeg / ffprobe for scaling, grading, mixing, and verification, targeting about -16 to -14 LUFS and true peaks no higher than -1 dBFS.
Boundaries
It assumes local ffmpeg, a transcription route that can produce timestamps, and image generation for the cover. Missing optional plugins should not block the final MP4. If the script or narration is missing and no safe fallback exists, the workflow should stop and request the asset. Sources should retain source pages, rights notes, and content hashes, not just filenames.
Use Cases
- With a Chinese script and voiced audio ready, produce a vertical 4K MP4, cover frame 0, and QA reports.
- Align narration ASR timestamps to the original script, generate non-overlapping one-line captions, and verify full-text exactness.
- Plan shots with real-person video, prioritize named speeches or interviews, and avoid looping, duplicates, and filename-only checks.
- Generate a black-gold 3:4 cover, inspect faces, text, and perspective, then embed it as the final video frame 0.
Best For
- Video creators who have a Chinese script and narration and need editing, captions, cover, and delivery validation.
- Content operators who need publishable vertical MP4 output with a source manifest and QA reports.
- Editing engineers who want a repeatable `ffmpeg` and transcription workflow for `4:3` cropping and mixing.
- Self-media editors who need real-person assets, license records, and a synced Jianying/CapCut draft.
Related Skills
Generate an interactive HTML product detail page from images and copy, with multi-product carousel, inline text editing, design controls, and PDF/JPG/PSD export support.
Enter a material name to get preview and download links for free, commercially usable video assets, with paid platform options, pricing, and search suggestions when results are limited.
A local CLI for Dreamina image and video generation, including credit checks, async submission, result queries, and task history review.
An AI-guided workflow for Chinese video dubbing and subtitles that configures iFlytek or edge-tts, splits audio, generates SRT, and composes the final video.