Three-Minute Book Reading Video Generator
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @user_25bf452a/book-video-generator.
About this skill
The Problem
Turning a book into a short video often gets stuck on three parts: copy that sounds natural, stable image generation, and audio-subtitle alignment. book-video-generator automates that pipeline: it takes book_name and author_name, then outputs a roughly three-minute MP4 explainer video, reducing manual scripting, image retries, and per-line caption work.
Core Workflow
The skill splits the job into five stages:
- Copy generation: it uses the current platform's web search to collect the book's intro, publication year, and common readings, then asks the LLM for a
700-1000word script. - Scenes and chapters: the script is split into
8-50scenes, each with subtitle text, visual description, and image prompt; four chapter titles are also generated for the top progress bar. - Asset generation: each scene gets a
1024x768flat illustration viaImageGen, Volcengine Jimeng,Gemini, orAgnes; subtitles are spoken with Volcengine TTS, falling back toedge-ttswithout credentials. - Composition:
compose_video.pyusesffmpegto combine images, audio,ASSsubtitles, cover frame, BGM, and transition sounds into a1920x1080video, with keyword highlighting and subtitle entrance/exit animation.
Boundaries
- Image and TTS generation depend on external models; CLI environments need the matching
API key, MCP, or script setup. - Volcengine TTS word timestamps are estimated from audio duration, while
edge-ttsnative timestamps are usually more precise. - BGM and page-turn audio may not be bundled; the video still renders without them if the files are missing.
- It is aimed at fixed-style book explainer shorts, not real-person narration, complex camera work, or multilingual dubbing.
Use Cases
- A book-channel creator turns one book into a three-minute explainer, needing script, scenes, images, TTS, and captions in one run.
- A content test team checks the same book under different image APIs, switching Volcengine, Gemini, or Agnes to regenerate scene art.
- A video editor breaks an LLM narration script into single-line ASS subtitles, synced to TTS timestamps with keyword highlights and animation.
- A cross-platform agent developer reuses the workflow in Codex CLI or TRAE Work, calling external image/TTS APIs to output an MP4 short.
Best For
- Creators running book-review channels who want fewer manual scenes, image retries, and per-line captioning tasks.
- Platform engineers validating agent toolchains who need to compare image generation, TTS, and ffmpeg composition across platforms.
- Content operators producing book explainer shorts who want a publishable MP4 from only book title and author.
- Automation engineers using WorkBuddy, OpenClaw, or TRAE Work who want to convert the Coze workflow into a portable skill.
Related Skills
Convert Chinese or natural-language requests into paste-ready English image prompts for ChatGPT's web UI, covering generation, editing, multi-image references, and exact text without performing image generation.
A creative AI image workflow for style transfer, reference-based creation, scene replacement, series expansion, material conversion, era shifts, composition rework, lineart conversion, and cartoonization.
Translates star-inspired football memories and fan resonance into compliant original IP poster concepts while avoiding real names, likenesses, official badges, and event logos.
Extracts real web page colors, typography, components, layouts, and interaction styles with a Playwright script to produce a verifiable design guideline.