PDF Office Processing Guide
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md and install @user_f39a06e7/pdf-official123.
About this skill
What problem it solves
PDF automation often gets stuck beyond simple merge/split: text extraction and table parsing from text-based PDFs, OCR for scanned pages, plus generating simple PDFs, adding watermarks, and setting passwords all require different tools. pdf-official2 maps these common operations to concrete tool paths so you do not have to guess across libraries.
How it works
It organizes workflows into Python libraries and command-line tools:
- pypdf for basic merge, split, metadata, and rotation;
- pdfplumber for layout-aware text and table extraction;
- reportlab for creating single- or multi-page PDFs;
- qpdf, pdftotext, and pdftk for quick merging, text conversion, or page handling;
- Scanned PDFs typically go through image conversion before pytesseract OCR.
A useful sequence is to identify the PDF type first: use pdfplumber for text-based files, preprocess scanned pages, and use reportlab Canvas or Platypus for generated reports.
Boundaries
It is aimed at scriptable, repetitive PDF tasks, not professional typesetting or complex document workflows. PDF form filling is directed to forms.md, and advanced usage/troubleshooting to reference.md. Real availability depends on installed Python packages and system tools; output quality is affected by the source PDF structure, table borders, and scan clarity.
Use Cases
- Split contract PDFs by page, then add page numbers and metadata to each file.
- Extract text and tables from financial reports and export CSV files for analysis.
- Merge project briefs into one PDF, then add a company watermark and password.
- Convert scanned pages to images, OCR the text, and assemble searchable output.
Best For
- Operations or finance staff who batch split/merge invoices, contracts, and reports
- Analysts who extract PDF report text and tables into structured data
- Product operators who generate simple spec PDFs with Python
- Legal assistants who convert scanned contracts to searchable text and add watermarks
Related Skills
Turns files, references, or chat context into styled single-page HTML reports with built-in templates and preset themes.
An engineering-oriented email automation solution for batch sending, Jinja2 templates, attachments, scheduled sending, inbox monitoring, and rule-based auto-reply.
An engineering workflow for creating, editing, reviewing, analyzing, and image-converting .docx files using pandoc, docx-js, OOXML, and LibreOffice.
Pre-submission scanner for Word/PDF blind-bid files that checks margins, fonts, page numbers, and metadata against tender requirements and flags residual bidder identities.