Xiangyun Open Platform Document and Table Recognition
Paste the following prompt into your AI chat to install this skill:
Please install @user_6b190ef3/doc-ocr-xy using https://skillhub.cn/install/skillhub.md.
About this skill
Problem Addressed
Documents such as scanned tables, invoices, and forms often exist as PDF, OFD, or image files, making manual transcription slow and error-prone. Xiangyun Open Platform Document and Table Recognition is useful for batch OCR of unstructured documents, producing processable text for downstream parsing, storage, or table extraction.
How the Skill Works
The skill relies on Xiangyun OCR credentials netocr_key and netocr_secret, which must be configured before first use. A typical flow is:
- Prepare input files: supports .pdf, .ofd, and image formats such as .jpg, .png, .bmp, .tif, .webp;
- Select a language code: for example, Simplified Chinese printed text uses 0, English uses 2, Japanese uses 11, and handwritten options map to other codes;
- Run batch recognition: files are sent to Xiangyun OCR, billed per call, and returned for further processing.
It fits engineering workflows that need multilingual documents normalized into text, especially scanned PDF/OFD files, receipts, and table screenshots.
Boundaries and Notes
The materials do not promise fully structured table output. If the goal is direct row and column extraction, treat OCR text as an intermediate result and add table parsing logic. Images should be clear, preferably larger than 500px in width and height; single files should stay under 10MB. The config file is stored as config.json in the skill directory. Credentials are sensitive and should not be committed to repositories or written to logs. Xiangyun OCR is billed per call, so validate cost and recognition quality on a small sample before running large batches.
Use Cases
- Batch-recognize table text in scanned PDF or OFD files as input for field extraction
- Convert printed Japanese, Korean, or Western-language documents from images into text for unified ingestion
- After preprocessing invoice and form images, call OCR to obtain text for rule-based validation
- Recognize text in multilingual scanned contracts before manual or programmatic review of key clauses
Best For
- Document processing engineers who need to convert scanned PDF or OFD files into text for field extraction
- Operations staff handling multilingual invoices and forms who need image text turned into validatable data
- Backend developers integrating third-party OCR and managing credentials plus batch jobs
- Compliance reviewers who need to recognize foreign-language contracts or tables before manual review
Related Skills
Organizes files by extension into subfolders like Documents, Code, and Archives, then outputs a report.
Extract tables, formulas, charts, and layout from invoices, reports, papers, and multi-column documents.
Generates a multi-sheet Excel report containing only structured data tables from byteplan-analysis results.
Automatically sort directory files into type-based folders, with dry-run preview, reports, and JSON custom rules.