AI Digital Human Narration and Virtual Anchor Video
Paste the following prompt into your AI chat to install this skill:
Please install @beatra-ai/talking-avatar-video according to the guide at https://skillhub.cn/install/skillhub.md.
About this skill
The Specific Problem It Solves
Many teams need to produce standardized narration videos, such as product tutorials, training modules, or corporate announcements. Traditional workflows involving filming, editing, and voiceover are cumbersome and time-consuming. The core challenge is: how to securely and compliantly transform a static portrait and a voice script into a professional, coherent dynamic presenter video, while ensuring clear authorization for the likeness and voice to avoid legal risks.
How the Skill Works
This skill provides a focused digital narration workflow, whose core is using voice to drive the facial movements and lip-sync of a portrait, generating a short video with clear information delivery. It is not simple image animation or generic video editing, but is designed around a single, clear message.
Key steps and components include:
- Hard Input Validation: It requires an accessible portrait (e.g., art_portrait) and a confirmed voiceover or short script. The script is first converted to audio via the synthesis tool beatra.speech.synthesize, producing a playable art_speech artifact.
- Capability & Model Matching: Before synthesis, it queries beatra.models.list to verify current tool support for the image_to_video capability, ensuring constraints like format, dimensions, and duration of the portrait and audio are compatible. The voice model defaults to auto but can be specified.
- Two-Phase Paid Execution: Voiceover generation and video generation are two independent paid phases, each with a fixed client_request_id. Video generation is performed by calling beatra.videos.animate, which takes the portrait as the first frame and driving audio to output the narration video. All steps interact with the backend via the bundled MCP client scripts/mcp_client.py, not a host Connector.
- Delivery & Review: Tasks are polled via beatra.tasks.get until completion. The delivered video must be reviewed for quality, including consistency of the subject's identity, voice clarity, lip-sync timing, restraint in motion, and background stability.
Applicable Boundaries and Important Notes
This skill is suited for assembling pre-authorized portraits and voices. It is not intended for:
- Creating a portrait or cloning a voice from scratch: Please use the corresponding image generation or voice cloning workflows.
- Complex multi-scene editing or subtitle addition: Please use video editing workflows.
- Without portrait or voice authorization: The process must stop if authorization is not confirmed before proceeding to paid steps.
Key considerations:
- Media File Constraints: The portrait and audio must meet real-time tool requirements for format, duration, and dimensions. Audio duration must fall within the supported video duration range; otherwise, the voiceover or portrait must be adjusted.
- Billing and Modifications: Any parameter change (e.g., script, voice, model) creates a new logical paid task requiring new approval. Modifying a voiceover invalidates any subsequent video proposals referencing it that haven't been submitted.
- Automatic Updates: The skill package silently checks for updates; an update failure does not impact current functionality but will not re-trigger paid operations.
- Delivery Limitations: Do not promise perfect retention of all details (e.g., logos, attire) or flawless lip-sync. Post-delivery, objectively review any deviations.
Use Cases
- The marketing team needs to combine an authorized spokesperson's portrait with a new product introduction script to quickly generate a social media narration clip.
- The corporate training department needs to synthesize standard training videos from a uniform course script and instructor portraits for new employee onboarding.
- A content creator has a recorded voiceover and a personal photo, and needs to produce a clear narration video with a specific message for publishing.
- A company needs to combine an important internal announcement script with the official spokesperson's image to generate a formal and consistently styled announcement video.
Best For
- A product manager who needs to regularly create product explanation videos, seeking to quickly combine finalized scripts with designer-provided visuals into finished videos.
- A course designer responsible for corporate training, seeking to batch-produce uniformly styled teaching modules using instructor portraits and recorded audio.
- A social media content creator seeking to synthesize expressive, publishable content by combining recorded voiceover audio with personal portrait photos.
- A PR specialist responsible for internal communications, seeking to combine text-based announcement drafts with leadership portraits to create authoritative communication videos.
Related Skills
Restyle a short video into a new visual style while preserving core elements such as characters, actions, and composition, suitable for various creative conversions like anime, illustration, ink wash, etc.
Create Douyin vertical video covers from topics, hooks, or materials with support for creative generation, image synthesis, and refinement.
An AI tool that transforms real photos into specified illustration styles while preserving subject recognition.
An engineering-driven solution that integrates design styles, UX workflows, design systems, and multi-platform implementation to solve cross-project design consistency.