Vidu Video and Image Generation
Paste the following prompt into your AI chat to install this skill:
Follow https://skillhub.cn/install/skillhub.md to install @user_1573f5dc/vidu-skills.
About this skill
Problem It Addresses
Image and video generation workflows often require manual handling of async API tasks, image uploads, reference subjects, and error triage. This skill wraps Vidu's vidu-cli in a CLI-first execution model: commands use flags rather than raw JSON bodies.
How It Works
Core capabilities include:
- Text-to-image / text-to-video: generate media from text prompts
- Image-to-video / start-end frame video: constrain video output with one image or start/end frames
- Reference-based image/video: combine 1–7 images or materials with a required text prompt
- Reference management: create, list, and search elements
- Async task query: submit a task, then poll with task get or task sse
The workflow is: submit, store task_id and trace_id, query until terminal state, then return media_urls or exact error fields. Both success and failure use a stdout JSON contract, which makes scripting straightforward.
Boundaries
It requires vidu-cli >= 0.2.0 and VIDU_TOKEN; it cannot run if the CLI is unavailable. Prompts, images, and task settings are sent to Vidu's API, so privacy and IP requirements must be confirmed. Use the bundled references/parameters.md for task types, model versions, and exact parameter rules.
Use Cases
- Generate product images from short copy for ad draft review.
- Turn a poster frame into a 3-second motion ad and poll for the final URL.
- Create a character reference and reference its name in the video prompt.
- Debug failures by reporting error.type, code, and message verbatim.
Best For
- Marketing editors who need to turn copy into image or video drafts
- Media planners maintaining character consistency through reference elements
- Automation engineers scripting async task polling and media URL retrieval
- Support engineers diagnosing vidu-cli failures from exact error fields
Related Skills
Uses step-by-step choices to confirm business, palette, and layout, then exports an editable .drawio architecture diagram.
An agent skill for the Miaoyin AI music REST API that supports SUNO/Mureka-based song generation, continuation, cover, video, WAV, and stem splitting.
Distills text, files, or images into structured knowledge and generates knowledge-card prompts for drawing tools, supporting 3:4 portrait or 16:9 landscape outputs.
Capture full-page, viewport, or element screenshots with Playwright while handling lazy loading, internal scroll, mobile layouts, and Linux CJK fonts.