AI Agent Hub
Back to skills
PDF Office Processing Guide icon

PDF Office Processing Guide

Office Efficiency Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md and install @user_f39a06e7/pdf-official123.

About this skill

What problem it solves

PDF automation often gets stuck beyond simple merge/split: text extraction and table parsing from text-based PDFs, OCR for scanned pages, plus generating simple PDFs, adding watermarks, and setting passwords all require different tools. pdf-official2 maps these common operations to concrete tool paths so you do not have to guess across libraries.

How it works

It organizes workflows into Python libraries and command-line tools:
- pypdf for basic merge, split, metadata, and rotation;
- pdfplumber for layout-aware text and table extraction;
- reportlab for creating single- or multi-page PDFs;
- qpdf, pdftotext, and pdftk for quick merging, text conversion, or page handling;
- Scanned PDFs typically go through image conversion before pytesseract OCR.

A useful sequence is to identify the PDF type first: use pdfplumber for text-based files, preprocess scanned pages, and use reportlab Canvas or Platypus for generated reports.

Boundaries

It is aimed at scriptable, repetitive PDF tasks, not professional typesetting or complex document workflows. PDF form filling is directed to forms.md, and advanced usage/troubleshooting to reference.md. Real availability depends on installed Python packages and system tools; output quality is affected by the source PDF structure, table borders, and scan clarity.

Use Cases

  • Split contract PDFs by page, then add page numbers and metadata to each file.
  • Extract text and tables from financial reports and export CSV files for analysis.
  • Merge project briefs into one PDF, then add a company watermark and password.
  • Convert scanned pages to images, OCR the text, and assemble searchable output.

Best For

  • Operations or finance staff who batch split/merge invoices, contracts, and reports
  • Analysts who extract PDF report text and tables into structured data
  • Product operators who generate simple spec PDFs with Python
  • Legal assistants who convert scanned contracts to searchable text and add watermarks