AI Agent Hub
Back to skills
📁

PDF Toolkit

Office Efficiency Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @org-02qudk26/pdf-zh according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

When PDFs enter an automated pipeline, the difficult part is usually not opening a file, but reliably extracting text and tables, merging or splitting documents, generating new PDFs, and writing form fields. PDF Toolkit focuses on this kind of programmatic document processing and provides a clear path from Python libraries to command-line tools. It organizes common PDF operations into an actionable toolkit, useful for data cleaning, report generation, and contract organization.

How It Works

  • Use pypdf for basic operations: merging, splitting, metadata reading, and page rotation.
  • Use pdfplumber to extract layout-aware text and tables, which is useful for converting PDF content into structured data.
  • Use reportlab with Canvas or Platypus to create PDFs, including multi-page documents.
  • Use pdftotext, qpdf, and pdftk for quick merging, conversion, or batch processing from the terminal.
  • Fill forms with pdf-lib or pypdf, following forms.md for field mapping and write operations.
  • For complex flows, consult reference.md for pypdfium2, JavaScript libraries, and troubleshooting guidance.

Boundaries

This skill is well suited to repeatable PDF operations in scripts or workflows. If the task requires OCR for scanned PDFs, the content usually needs to be converted to images and recognized first. If the focus is advanced rendering, complex layout understanding, or interactive form handling, review the relevant reference documents instead of relying only on basic merge and split utilities.

Use Cases

  • When processing contract batches, use `qpdf` or `pypdf` to merge, split by page, and archive PDFs.
  • When analyzing research PDFs, use `pdfplumber` to extract text and tables, then clean them into structured data.
  • When generating performance reports, use `reportlab` with `Canvas` or `Platypus` to output multi-page PDFs.
  • When organizing scanned materials, use `pytesseract` to convert image-based content into searchable text first.

Best For

  • Backend engineers: integrate PDF parsing, splitting, and form filling into automation pipelines.
  • Data analysts: extract text and tables from research or financial PDFs for downstream cleaning.
  • Report developers: generate multi-page PDF reports programmatically with `reportlab`.
  • DevOps engineers: run command-line PDF batch processing with `qpdf` and `pdftotext`.