Document Reader
Paste the following prompt into your AI chat to install this skill:
Follow https://skillhub.cn/install/skillhub.md to install @user_053d27d0/user-053d27d0-document-reader.
About this skill
Problem to solve
When an AI needs to inspect attached files or documents inside archives, formats are fragmented: PDF needs text extraction, xlsx needs sheet-level expansion, pptx needs slide-level chunking, and zip, rar, or 7z must be listed before a target file can be selected. Manual preprocessing interrupts context and is awkward for batch review.
How it works
Document Reader turns office documents and archives into a “list → select → output text” workflow:
- Office documents: supports
pdf,docx,xlsx,pptx,rtf,odt,html/htm, and plain-text files such astxt,md,json,xml,py, andjs. - Archives: supports
zip,tar,tgz,tar.bz2,rar, and7z; it can first list contents and then read a specified document. - Output modes: human-readable text or a
JSONinterface; file matching is case-insensitive and falls back to fuzzy matching when an exact match is unavailable. - Structure handling:
xlsxoutput is organized by sheet as table text, whilepptxis chunked by slide for easier section-level analysis.
Boundaries and notes
It is designed for content extraction rather than document editing: it suits AI analysis, quick attachment review, data extraction, and preprocessing, but not modifying source files, preserving complex layouts, or reconstructing formulas. pdf extraction depends on poppler-utils, so scanned or encrypted PDFs may have limited fidelity. Archive support lists and reads internal documents without automatically extracting them to disk.
Use Cases
- List files inside a vendor zip attachment, then extract the quote xlsx and contract PDF text for AI comparison.
- Read meeting notes from multiple pptx and docx files as paragraphs, output JSON for an LLM to summarize decisions and action items.
- Match an HTML report and txt log by fuzzy filename in a tar.gz project package, then extract body text for debugging.
- Read metric sheets from an Excel workbook sheet by sheet, convert them to table text, and ask an AI to reconcile definitions and differences.
Best For
- Sales or procurement staff reviewing vendor quote packages need to locate PDF and xlsx files inside a zip and extract text for comparison.
- PMs or assistants organizing meeting minutes need to turn pptx/docx paragraphs into AI-ready summaries of decisions and action items.
- Ops or data engineers handling project archives need to list tar.gz or 7z contents and read HTML reports and txt logs for debugging.
- Algorithm engineers preprocessing document data need to read office documents in batches and emit JSON or text for model consumption.
Related Skills
Batch-import team weekly reports, aggregate progress, plans, issues, and support by project, flag delivery or resource risks, and generate structured department reports.
Generates Kingdee Cloud ERP startup, implementation, and acceptance documents from Word templates with delivery guidance.
Extract decisions and action items from meeting notes, emails, or chat logs, then track owners, deadlines, and follow-ups with lightweight reminders and no direct tool integrations.
Exports PPT slides in groups of N into high-resolution vertically merged images, auto-named by extracted student names for batch report export.