AI Agent Hub
Back to skills
📁

Document Reader

Office Efficiency Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Follow https://skillhub.cn/install/skillhub.md to install @user_053d27d0/user-053d27d0-document-reader.

About this skill

Problem to solve

When an AI needs to inspect attached files or documents inside archives, formats are fragmented: PDF needs text extraction, xlsx needs sheet-level expansion, pptx needs slide-level chunking, and zip, rar, or 7z must be listed before a target file can be selected. Manual preprocessing interrupts context and is awkward for batch review.

How it works

Document Reader turns office documents and archives into a “list → select → output text” workflow:

  • Office documents: supports pdf, docx, xlsx, pptx, rtf, odt, html/htm, and plain-text files such as txt, md, json, xml, py, and js.
  • Archives: supports zip, tar, tgz, tar.bz2, rar, and 7z; it can first list contents and then read a specified document.
  • Output modes: human-readable text or a JSON interface; file matching is case-insensitive and falls back to fuzzy matching when an exact match is unavailable.
  • Structure handling: xlsx output is organized by sheet as table text, while pptx is chunked by slide for easier section-level analysis.

Boundaries and notes

It is designed for content extraction rather than document editing: it suits AI analysis, quick attachment review, data extraction, and preprocessing, but not modifying source files, preserving complex layouts, or reconstructing formulas. pdf extraction depends on poppler-utils, so scanned or encrypted PDFs may have limited fidelity. Archive support lists and reads internal documents without automatically extracting them to disk.

Use Cases

  • List files inside a vendor zip attachment, then extract the quote xlsx and contract PDF text for AI comparison.
  • Read meeting notes from multiple pptx and docx files as paragraphs, output JSON for an LLM to summarize decisions and action items.
  • Match an HTML report and txt log by fuzzy filename in a tar.gz project package, then extract body text for debugging.
  • Read metric sheets from an Excel workbook sheet by sheet, convert them to table text, and ask an AI to reconcile definitions and differences.

Best For

  • Sales or procurement staff reviewing vendor quote packages need to locate PDF and xlsx files inside a zip and extract text for comparison.
  • PMs or assistants organizing meeting minutes need to turn pptx/docx paragraphs into AI-ready summaries of decisions and action items.
  • Ops or data engineers handling project archives need to list tar.gz or 7z contents and read HTML reports and txt logs for debugging.
  • Algorithm engineers preprocessing document data need to read office documents in batches and emit JSON or text for model consumption.