PaddleOCR Document Parsing
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md and install @user_15292d5a/yjkj-paddleocr-doc-parsing
About this skill
Problem It Solves
Many office documents are not plain text: invoices, financial reports, spreadsheets, papers, and multi-column materials mix tables, formulas, charts, and layout relationships. Reading raw text often loses row and column structure, formula meaning, chart positions, or fails on curved, folded, or mis-oriented scans.
How It Works
The skill targets PaddleOCR document parsing and can start from a URL or a local file. It focuses on complex layout extraction: tables, math formulas, charts, multi-column layouts, and overall structure analysis. For flat screenshots or properly scanned pages, preprocessing can be disabled for faster results. For curved or folded photos, perspective distortion, or rotated inputs, preprocessing should remain enabled so the image is corrected before parsing.
The output should show complete extracted content by default, truncating only around 10,000 characters. Multi-page content can be summarized, but full output is provided when explicitly requested.
Boundaries
It is better suited for structured document reading than simple OCR text capture. Missing or invalid PADDLEOCR_ACCESS_TOKEN values, API quota limits, blank pages, no extractable text, and low-quality scans should be reported as specific issues rather than failing silently.
Use Cases
- Parse invoice and financial report screenshots into structured tables for reconciliation.
- Extract formulas from academic papers and keep them aligned with surrounding text.
- Read multi-column brochures or newsletters and restore logical text order.
- Capture chart-related text from report pages for later manual data review.
Best For
- Finance staff reconciling invoices and reports who need itemized table extraction.
- Researchers processing papers who need formulas kept with context.
- Analysts archiving competitor materials who need multi-column text converted.
- Integration engineers building parsing pipelines that handle URLs, local scans, and specific errors.
Related Skills
Organizes files by extension into subfolders like Documents, Code, and Archives, then outputs a report.
Generates a multi-sheet Excel report containing only structured data tables from byteplan-analysis results.
Automatically sort directory files into type-based folders, with dry-run preview, reports, and JSON custom rules.
A six-stage pipeline for turning .docx templates and source materials into enterprise documents with preserved styles, mapped TOCs, Mermaid diagrams, and token-budget controls.