Deli File OCR Parser
Paste the following prompt into your AI chat to install this skill:
Please install @user_f26b93e2/deli-ocr-file-parse according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem It Solves
In contract review, evidence organization, and invoice extraction workflows, agents often encounter scanned PDFs, images, OFD files, and Office tables. Native parsing may miss pages, produce garbled text, or break table structure, making downstream analysis unreliable. This skill normalizes file parsing into reusable Markdown or .txt output, with Deli OCR as a fallback when native capabilities are insufficient.
How It Works and Where It Applies
The skill first prefers the current agent or platform parsing capabilities:
- PDF: try platform parsing, pdfplumber, PyMuPDF, or similar local tools.
- Images/scans: try multimodal recognition or built-in OCR.
- Office/spreadsheets: try python-docx, openpyxl, CSV, or JSON parsing.
- Text/HTML/Markdown: read directly without OCR.
It calls deli-cli only when native parsing returns empty results, missing pages, garbled text, unreadable tables or invoice fields, or when the user explicitly asks for Deli OCR. Before invocation, it completes the CLI prerequisite check and follows the command, parameters, and RUN information returned at runtime rather than assuming fixed commands. It supports common .pdf, .docx, .xlsx, .png, .ofd, and .html formats, with actual capabilities determined by the current command response.
Caveats
After parsing, it should report the source filename, output path, whether raw response was saved, and any uncertainty such as missing pages, garbled text, table misalignment, stamps, or handwritten content. It should not call OCR automatically just because an API key is configured. Large parsed text should be written to files before being passed to other tools, and original files must be preserved. Results involving amounts, dates, case numbers, invoice numbers, or bank account numbers require manual review. For sensitive materials, only output the fragments needed for the task.
Use Cases
- Legal teams convert scanned PDF evidence into readable Markdown to review dates, amounts, and clauses item by item.
- Compliance staff parse OFD or image contracts into text to check terms, case numbers, and stamp readability.
- Finance reviewers extract invoice images or PDFs into Markdown to capture invoice numbers, amounts, and dates for verification.
- Engineers batch-convert Office docs or HTML pages into Markdown for downstream documentation analysis.
Best For
- Lawyers organizing evidence who need scanned PDFs converted into readable text for clause-by-clause review.
- Finance or operations reviewers who need invoice images parsed into invoice numbers, amounts, and dates.
- Legal staff maintaining contract libraries who need OFD or image contracts converted to Markdown for review.
- Agent workflow engineers who need an OCR fallback when native file parsing fails.
Related Skills
Organizes files by extension into subfolders like Documents, Code, and Archives, then outputs a report.
Extract tables, formulas, charts, and layout from invoices, reports, papers, and multi-column documents.
Generates a multi-sheet Excel report containing only structured data tables from byteplan-analysis results.
Automatically sort directory files into type-based folders, with dry-run preview, reports, and JSON custom rules.