PDF Processing Guide
Paste the following prompt into your AI chat to install this skill:
Follow https://skillhub.cn/install/skillhub.md and install @user_f39a06e7/docx12345544.
About this skill
Problem Addressed
PDF files in office automation are often unstructured: merging, splitting, extracting text or tables, generating simple reports, adding watermarks, and protecting files are not well suited to manual workflows or a single GUI. This skill provides an engineer-oriented PDF processing path based on Python and command-line tools, focusing on repeatable, scriptable document tasks.
How It Works
- Basic document operations: use
pypdffor merging, splitting, extracting metadata, and rotating pages. - Content extraction: use
pdfplumberto extract layout-aware text and tables, which helps convert PDF information into structured data. - Document creation: use
reportlabwithCanvasorPlatypusto generate single-page or multi-page PDFs. - Command-line support: use tools such as
pdftotextandqpdffor batch text conversion, merging, and related operations. - Scanned documents and common tasks: convert scans to images and use
pytesseractfor OCR, with coverage for watermarks, password protection, and similar scenarios. - Forms and advanced use: PDF form filling should follow
forms.md, while JavaScript libraries, detailed examples, and troubleshooting content are inreference.md.
Boundaries
This skill is suitable for automation scripts, batch office document processing, and simple PDF generation. It does not replace professional layout tools and does not guarantee results on complex document recognition. For forms, advanced pypdfium2, or pdf-lib usage, follow the corresponding referenced documentation.
Use Cases
- Split a batch of monthly PDF reports by page into separate files for archiving.
- Extract text and tables from PDF contracts or reports before importing them into systems.
- Generate multi-page PDF deliverables in bulk from fixed templates with reportlab.
- Convert scanned PDFs to images and run OCR to produce searchable text.
Best For
- Backend engineers building automation scripts who need stable PDF merging, splitting, and metadata extraction.
- Data analysts handling contracts and reports who need structured data from PDF tables.
- Ops engineers producing deliverables who need bulk multi-page PDF generation.
- Technical support staff processing scans who need to convert scanned PDFs to images for OCR.
Related Skills
Organizes files by extension into subfolders like Documents, Code, and Archives, then outputs a report.
Extract tables, formulas, charts, and layout from invoices, reports, papers, and multi-column documents.
Generates a multi-sheet Excel report containing only structured data tables from byteplan-analysis results.
Automatically sort directory files into type-based folders, with dry-run preview, reports, and JSON custom rules.