AI Agent Hub
Back to skills
Course Video Digital Human Instructor icon

Course Video Digital Human Instructor

Education Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @beatra-ai/course-video-studio into your AI assistant.

About this skill

Problem

Making instructor-led course videos is often blocked less by script quality than by workflow: a licensed instructor portrait, final lesson scripts, voice rights, and video models that accept specific image and audio formats must all line up. This skill targets that gap: it turns approved scripts and one authorized instructor portrait into digital-human talking-head course videos, delivered in lesson order.

How it works

  • Input freeze: Check the portrait's MIME type, width, height, aspect ratio, byte size, and alpha channel, then create a lesson ledger. Scripts, likeness and voice rights, and pronunciation tables must be explicit; if repeated names, terms, or abbreviations appear without a pronunciation table, the workflow stops to collect readings.
  • Voice selection: Use a catalog voice or, after consent and balance confirmation, clone a voice. Cloning goes through a confirmation card before beatra.voices.clone, then freezes the returned voice_id.
  • TTS and video compatibility: Query TTS and image_to_video capabilities, compare the real audio mime_type, duration, and size against accepted constraints, and split the pilot lesson at sentence or paragraph boundaries into supported segments.
  • Lesson delivery: Each audio segment must be confirmed before video generation. The video length is chosen as the smallest supported integer duration that is not shorter than the real audio, and final outputs are delivered by lesson and segment order with real size, duration, and usage.

Boundaries

It does not replace audio-only course packages, single talking-avatar clips, or slide parsing. Slide content only matters after the user extracts speakable text into the script. Likeness, voice, and billing authorization must be confirmed first; if formats are incompatible or balance is insufficient, the workflow stops before paid calls instead of silently changing formats or guessing outcomes.

Use Cases

  • Online course teams turn finalized multi-lesson scripts and one licensed instructor portrait into voiced talking-head course videos.
  • Training leads split a term-dense pilot lesson into supported segments, verify pronunciation and audio duration before generating video.
  • Education operations prepare weekly reviewable digital-instructor clips, deliver them in lesson order with real size, duration, and usage.
  • Independent instructors confirm voice-cloning consent and balance, then turn one final script into a video segment with confirmed audio.

Best For

  • Online course producers who need to convert a stable instructor portrait and final scripts into deliverable talking-head lesson videos.
  • Corporate training leads who must confirm terminology pronunciation, voice rights, and video segmentation before producing course clips.
  • Education product engineers who need to control TTS and image_to_video calls based on real audio duration and model format constraints.
  • Independent instructors who want video generation to start only after each audio segment is confirmed, then review terminology and lip timing.