AI Agent Hub
Back to skills
PDF Processing Guide icon

PDF Processing Guide

Office Efficiency Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Follow https://skillhub.cn/install/skillhub.md and install @user_f39a06e7/docx12345544.

About this skill

Problem Addressed

PDF files in office automation are often unstructured: merging, splitting, extracting text or tables, generating simple reports, adding watermarks, and protecting files are not well suited to manual workflows or a single GUI. This skill provides an engineer-oriented PDF processing path based on Python and command-line tools, focusing on repeatable, scriptable document tasks.

How It Works

  • Basic document operations: use pypdf for merging, splitting, extracting metadata, and rotating pages.
  • Content extraction: use pdfplumber to extract layout-aware text and tables, which helps convert PDF information into structured data.
  • Document creation: use reportlab with Canvas or Platypus to generate single-page or multi-page PDFs.
  • Command-line support: use tools such as pdftotext and qpdf for batch text conversion, merging, and related operations.
  • Scanned documents and common tasks: convert scans to images and use pytesseract for OCR, with coverage for watermarks, password protection, and similar scenarios.
  • Forms and advanced use: PDF form filling should follow forms.md, while JavaScript libraries, detailed examples, and troubleshooting content are in reference.md.

Boundaries

This skill is suitable for automation scripts, batch office document processing, and simple PDF generation. It does not replace professional layout tools and does not guarantee results on complex document recognition. For forms, advanced pypdfium2, or pdf-lib usage, follow the corresponding referenced documentation.

Use Cases

  • Split a batch of monthly PDF reports by page into separate files for archiving.
  • Extract text and tables from PDF contracts or reports before importing them into systems.
  • Generate multi-page PDF deliverables in bulk from fixed templates with reportlab.
  • Convert scanned PDFs to images and run OCR to produce searchable text.

Best For

  • Backend engineers building automation scripts who need stable PDF merging, splitting, and metadata extraction.
  • Data analysts handling contracts and reports who need structured data from PDF tables.
  • Ops engineers producing deliverables who need bulk multi-page PDF generation.
  • Technical support staff processing scans who need to convert scanned PDFs to images for OCR.