Virtual Singer MV Production
Paste the following prompt into your AI chat to install this skill:
Please install @user_a018d531/virtual-singer-mv into your AI assistant following https://skillhub.cn/install/skillhub.md.
About this skill
Problem
Creating a virtual-singer MV usually requires aligning lyrics to beats, generating scene footage, keeping the character consistent, animating the singer, and compositing subtitles with the original track. This skill targets a concrete production case where the inputs include music audio, lyric text, a video script or storyboard, and a virtual singer three-view reference image. It organizes the workflow into staged, user-confirmable steps rather than relying on one opaque model call.
How It Works
The pipeline first analyzes audio with openai-whisper-api and scripts/analyze_audio.py, extracting BPM, beat timestamps, and duration to create lyrics_timeline.json and scene_timeline.json. It then generates scene clips with VideoGen from text or concept images, optionally supplementing material with media-downloader. Character images are produced from the three-view reference using ImageGen, then refined and upscaled with nano-banana-pro. Animation can use VideoGen image-to-video, video-fx stylized templates, or image-to-3D model generation. Finally, remotion-video-toolkit composites scenes, character animation, original audio, and lyric subtitles into a 1080P video, with FFmpeg as a fallback if rendering fails.
Boundaries
The skill is intended for scene-based MVs, not general music generation or lyrics-only creation. If the audio is mainly instrumental, Whisper may not produce usable lyrics, so supplied lyric text becomes important. Failures in VideoGen, 3D generation, or Remotion rendering can fall back to still images, template animation, or FFmpeg assembly. Character consistency depends on repeated key visual descriptors, and the default output is 1080P at 30 FPS.
Use Cases
- Create a 1080P MV with character animation and subtitles from music, lyrics, storyboard, and reference images.
- Align mainly instrumental audio and supplied lyrics to beats, output a lyric timeline, and composite captions.
- Generate stage or city-night scene clips from a script and supplement B-roll material as needed.
- Produce multi-expression singer images from three-view references, then add template animation or 3D rotation shots.
Best For
- Independent musicians who have a song, lyrics, storyboard, and reference images and want a publishable virtual-singer MV.
- Animation short-film creators who need to turn character three-view references into expression assets, template animations, and a timeline.
- AIGC video creators who want audio analysis, scene generation, and Remotion compositing into a 1080P MV.
- Virtual idol operators who need consistent character visuals and stage-performance videos with lyric captions.
Related Skills
Convert Chinese or natural-language requests into paste-ready English image prompts for ChatGPT's web UI, covering generation, editing, multi-image references, and exact text without performing image generation.
A creative AI image workflow for style transfer, reference-based creation, scene replacement, series expansion, material conversion, era shifts, composition rework, lineart conversion, and cartoonization.
Translates star-inspired football memories and fan resonance into compliant original IP poster concepts while avoiding real names, likenesses, official badges, and event logos.
Extracts real web page colors, typography, components, layouts, and interaction styles with a Playwright script to produce a verifiable design guideline.