AI Agent Hub
Back to skills
📁

PDF Reader Assistant

Office Efficiency Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_2dc8a2f8/pdf-reader-assistant.

About this skill

Problem

PDF analysis often gets stuck in three practical cases: readable text PDFs are easy, but scanned pages, dense tables, and figure annotations need extra handling; a single research report, paper, or contract needs TOC, summary, keywords, and table extraction; and multiple documents need side-by-side comparison instead of manual copying. PDF Reader Assistant targets these engineering-side reading tasks by turning extraction, structured analysis, and multi-PDF comparison into a reusable workflow.

How it works

  • Text extraction: prefers readable text, then estimates scanned content from the first 5 pages and can switch to OCR, trying pytesseract and then easyocr.
  • Structured analysis: outputs TOC, head/middle/tail summary, Top 15 keywords, Markdown tables, image annotations, and stats such as total words, Chinese characters, English words, and numeric occurrences.
  • Multi-PDF comparison: quickly summarizes several PDFs, then compares common keywords and key differences.
  • Batch processing: scans a directory for .pdf files and produces a summary table with filename, page count, word count, Top 5 keywords, and table count.
  • Large-file strategy: for documents over 50 pages, it first extracts the first 5 pages and TOC, then reads selected page ranges to avoid loading too much context at once.

Boundaries and notes

  • Encrypted PDFs require a password before reading.
  • Table extraction and OCR are optional enhancements; missing dependencies are skipped and noted.
  • Chinese scanned-page recognition requires simplified-Chinese OCR language support.
  • Extracted data is written to a temporary file and cleaned up after use.

Use Cases

  • Read a 50+ page research report by previewing the first 5 pages, then drilling into selected sections for key data.
  • Extract key clauses, tables, and statistics from a contract PDF to build a summary for review.
  • Compare three supplier PDF proposals by summarizing each one and listing common keywords and differences.
  • Scan a project folder and generate pages, word counts, Top 5 keywords, and table counts for every PDF.

Best For

  • Analysts who need to decide whether a long PDF is worth reading, using TOC and summary to locate key sections.
  • Legal staff handling scanned papers or contracts, needing body text, tables, and keywords when clean text is missing.
  • Operations staff comparing supplier or competitor PDFs, needing summaries, shared keywords, and key differences.
  • Docs engineers organizing a batch of folders, needing pages, word counts, Top 5 keywords, and table counts per PDF.