AIM Digital Human Video Generator
Paste the following prompt into your AI chat to install this skill:
Follow https://skillhub.cn/install/skillhub.md and install @user_34d5ae4c/aim-digital-human-video into your AI assistant.
About this skill
Problem It Solves
When a user already has a portrait image and wants it to “speak” from a specified audio file or script, the workflow often involves TTS, base64 encoding, gateway submission, status polling, and secret-key placement. This skill collapses those steps into a repeatable pipeline: provide an image and either audio or text, then receive a directly usable video URL.
Core Workflow
- Input: local path to a person image; audio source can be one of: existing
--audio-path, text plus voice sample--text/--voice-sample/--voice-sample-text, or text plus preset voice--voice-preset. - Generation: if text is provided, TTS is first called through the AEP gateway; then the image and audio are read, converted to base64, and submitted to
/video2, where the server returns a public TOS URL. - Polling:
ffprobereads the audio duration, and the skill adjusts HTTP HEAD polling frequency by stage to reduce unnecessary requests.
Boundaries and Notes
- The secret key is written only to the skill root
.envfile, under the keyaim-secret-key; a preflight check is required when it is missing. - A single polling run defaults to a 60-minute timeout, but timeout does not mean failure;
.task-history.jsonlis used for later re-checks. - Action-description
promptis not exposed by default; when no audio and no voice preset are provided, the “Chinese female” preset is used.
Use Cases
- Create a product-update video by driving a customer portrait with a script and Chinese female TTS.
- Turn an existing brand narration audio file into a talking-head video from a provided portrait.
- Clone a voice sample, then generate a promotional video from a script and a person image.
- Submit a generation task, poll for completion, and return the public video URL.
Best For
- Marketing content makers who need to turn a customer portrait into a talking video.
- Operations staff who have brand audio or scripts and need quick digital-human narration.
- Engineers who need to call the AEP gateway and poll the returned TOS video URL.
- Integration staff who configure aim-secret-key and rerun the digital-human generation script.
Related Skills
Uses step-by-step choices to confirm business, palette, and layout, then exports an editable .drawio architecture diagram.
An agent skill for the Miaoyin AI music REST API that supports SUNO/Mureka-based song generation, continuation, cover, video, WAV, and stem splitting.
Distills text, files, or images into structured knowledge and generates knowledge-card prompts for drawing tools, supporting 3:4 portrait or 16:9 landscape outputs.
Capture full-page, viewport, or element screenshots with Playwright while handling lazy loading, internal scroll, mobile layouts, and Linux CJK fonts.