AI Agent Hub
Back to skills
Omni Reader Multimodal Document Parsing icon

Omni Reader Multimodal Document Parsing

Office Efficiency Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @org-y1eh5fyg/cue-omni-reader according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

Multimodal inputs are often scattered across PDFs, Office files, images, web pages, audio, and video. When engineers need searchable text for agents, common issues include lost reading order, broken table structure, unreadable scans, and missing visual semantics in audio or video.

How It Works

omni-reader parses inputs into markdown or clean text. PDFs and Office documents preserve reading order and table structure; images and scans use OCR plus visual understanding; audio is transcribed with ASR; video combines ASR, keyframe visual analysis, and multimodal fusion. URLs can be parsed directly, while local files are processed through user-specified paths. Status, result reading, cancellation, and export actions are exposed through MCP tools, and backend services can also integrate via HTTP API.

Boundaries

Each run processes one file, with a single-file limit of roughly 256 MiB. Large files should be split, and video should be chunked into 15–30 minute segments. The default is no_store=true, so source files and results are not uploaded to the service; however, actual parsing depends on the remote service and may be affected by network, quota, peak hours, or service status. Unreachable addresses such as oss:// URLs should first be converted to signed HTTPS URLs.

Use Cases

  • Turn meeting recordings into searchable text minutes for later citation
  • Convert scanned forms and handwritten tables into structured text for project docs
  • Extract video subtitles and key visual moments into contextual technical summaries
  • Normalize web pages, tables, and document notes into one searchable text set

Best For

  • Project assistants who need meeting audio turned into searchable text minutes
  • Operations staff who process scanned contracts and handwritten tables
  • Training leads who compile video narration and subtitles into searchable materials
  • Engineers who feed document notes, web pages, and tables into agent knowledge bases