PDF Intelligence Suite
Paste the following prompt into your AI chat to install this skill:
Please install @user_9ebebf30/calenderv2222 according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem context
PDFs often act as isolated information sources in automation pipelines: the body may be flowing text or a scanned image, tables need to be normalized into CSV or Excel, and files may still require merging, splitting, watermarking, or metadata edits. Manual handling is costly and can easily disturb layout, page order, or permission-related details.
Capabilities and workflow
The skill is organized around PDF parsing and output:
- Text extraction: Uses
pdfplumberto obtain plain text or position-aware structured text, which helps locate fields. - Table recognition: Combines
camelot-pyandpdfplumberto extract PDF tables into structured data, withCSV/Excel-friendly output. - OCR recognition: Uses
pytesseractto handle scanned documents and image-based PDFs, supporting multilingual text extraction. - Format conversion: Uses
pdf2imageto convert PDFs into images and provides intermediate results for downstream workflows targeting Word or Excel. - Page and security operations: Uses
PyPDF2to merge, split, rotate, delete, encrypt, or decrypt pages, and can also add watermarks or digital signatures. - Metadata: Reads and modifies document properties to support archiving and traceability.
Boundaries and cautions
OCR requires an external Tesseract environment; complex layouts, cross-page tables, encrypted files, or low-quality scans may still affect recognition accuracy. Digital signatures and encryption are document-security operations, so file ownership and permission requirements should be clarified before use.
Use Cases
- When processing scanned procurement contracts, run OCR to extract key clauses and merge them into one searchable PDF.
- When organizing supplier quote PDFs, use table recognition to export price tables to CSV for reconciliation in Excel.
- Before archiving monthly financial reports, batch-delete draft pages, rotate landscape pages, and add confidentiality watermarks.
- Read bid document metadata and update author, title, and creation date to keep archive information consistent.
Best For
- Legal assistants handling contract archives who need to extract clauses from scans and add confidentiality watermarks.
- Financial analysts reconciling supplier data who need to turn PDF quote tables into CSV or Excel.
- Office automation engineers maintaining document libraries who need to merge, split, encrypt, and edit PDF metadata.
- Overseas operations staff handling multilingual scans who need OCR text extraction for later translation.
Related Skills
Organizes files by extension into subfolders like Documents, Code, and Archives, then outputs a report.
Extract tables, formulas, charts, and layout from invoices, reports, papers, and multi-column documents.
Generates a multi-sheet Excel report containing only structured data tables from byteplan-analysis results.
Automatically sort directory files into type-based folders, with dry-run preview, reports, and JSON custom rules.