AI Agent Hub
Back to skills
🎨

Three-Minute Book Reading Video Generator

Design & Media Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_25bf452a/book-video-generator.

About this skill

The Problem

Turning a book into a short video often gets stuck on three parts: copy that sounds natural, stable image generation, and audio-subtitle alignment. book-video-generator automates that pipeline: it takes book_name and author_name, then outputs a roughly three-minute MP4 explainer video, reducing manual scripting, image retries, and per-line caption work.

Core Workflow

The skill splits the job into five stages:

  • Copy generation: it uses the current platform's web search to collect the book's intro, publication year, and common readings, then asks the LLM for a 700-1000 word script.
  • Scenes and chapters: the script is split into 8-50 scenes, each with subtitle text, visual description, and image prompt; four chapter titles are also generated for the top progress bar.
  • Asset generation: each scene gets a 1024x768 flat illustration via ImageGen, Volcengine Jimeng, Gemini, or Agnes; subtitles are spoken with Volcengine TTS, falling back to edge-tts without credentials.
  • Composition: compose_video.py uses ffmpeg to combine images, audio, ASS subtitles, cover frame, BGM, and transition sounds into a 1920x1080 video, with keyword highlighting and subtitle entrance/exit animation.

Boundaries

  • Image and TTS generation depend on external models; CLI environments need the matching API key, MCP, or script setup.
  • Volcengine TTS word timestamps are estimated from audio duration, while edge-tts native timestamps are usually more precise.
  • BGM and page-turn audio may not be bundled; the video still renders without them if the files are missing.
  • It is aimed at fixed-style book explainer shorts, not real-person narration, complex camera work, or multilingual dubbing.

Use Cases

  • A book-channel creator turns one book into a three-minute explainer, needing script, scenes, images, TTS, and captions in one run.
  • A content test team checks the same book under different image APIs, switching Volcengine, Gemini, or Agnes to regenerate scene art.
  • A video editor breaks an LLM narration script into single-line ASS subtitles, synced to TTS timestamps with keyword highlights and animation.
  • A cross-platform agent developer reuses the workflow in Codex CLI or TRAE Work, calling external image/TTS APIs to output an MP4 short.

Best For

  • Creators running book-review channels who want fewer manual scenes, image retries, and per-line captioning tasks.
  • Platform engineers validating agent toolchains who need to compare image generation, TTS, and ffmpeg composition across platforms.
  • Content operators producing book explainer shorts who want a publishable MP4 from only book title and author.
  • Automation engineers using WorkBuddy, OpenClaw, or TRAE Work who want to convert the Coze workflow into a portable skill.