AI Agent Hub
Back to skills
Vidu Video, Image, and Audio Generation icon

Vidu Video, Image, and Audio Generation

Design & Media Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Follow https://skillhub.cn/install/skillhub.md to install @vidu/vidu-generation-2.

About this skill

Problem

Calling Vidu APIs directly requires juggling separate endpoints for video, image, and audio generation, plus model choice, duration, resolution, voice IDs, domain selection, and async polling. Developers who describe a task in natural language still need to decide whether to use text-to-video, image-to-video, or TTS, and must handle VIDU_API_KEY, generated URL expiry, and task status.

How the skill works

The skill maps natural-language requests to Vidu generation types: plain text maps to text-to-video or text-to-image; a first-frame image maps to image-to-video; multiple reference images map to reference-based video; two first-and-last frames map to transition video; dubs, narration, or voice cloning map to audio workflows.
- Model selection: recommends models such as viduq3-pro, viduq3-turbo, and q3-fast based on the scenario, then handles duration, resolution, and special aspect ratios.
- Audio synthesis: supports speed, volume, pitch, and emotion, and recommends Mandarin, English, or other voice IDs by content style.
- Execution: checks VIDU_API_KEY before calls, selects api.vidu.cn or api.vidu.com by user language, polls async tasks, and returns the available creations[0].url.

Boundaries and notes

The skill relies on a Vidu open-platform API key and does not handle key registration, billing, or asset rights. Generated URLs are usually valid for 24 hours, and cloned voices must be used within 7 days to persist. Long-form TTS, image size limits, model capability differences, and task failures still require checking API error messages.

Use Cases

  • Short-video production: turn a script and first frame into a 5-second audio-synced video for social posting.
  • Multi-character ads: upload character reference images to generate a 3-10 second multi-subject video and review cuts.
  • Voiceover: synthesize long tutorial text as a Chinese male announcer TTS and adjust speed and pauses.
  • Image work: create 1K/4K posters from reference images or perform local repaint with Vidu image models.

Best For

  • Operations staff who need to turn product scripts into social short videos quickly
  • Editors who need stable Chinese announcer voiceovers for tutorial copy
  • Engineers integrating Vidu APIs for multimodal generation workflows
  • Visual designers creating character reference sheets and multi-subject video drafts