AI Agent Hub
Back to skills
Xiangyun Open Platform Document and Table Recognition icon

Xiangyun Open Platform Document and Table Recognition

Office Efficiency Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please install @user_6b190ef3/doc-ocr-xy using https://skillhub.cn/install/skillhub.md.

About this skill

Problem Addressed

Documents such as scanned tables, invoices, and forms often exist as PDF, OFD, or image files, making manual transcription slow and error-prone. Xiangyun Open Platform Document and Table Recognition is useful for batch OCR of unstructured documents, producing processable text for downstream parsing, storage, or table extraction.

How the Skill Works

The skill relies on Xiangyun OCR credentials netocr_key and netocr_secret, which must be configured before first use. A typical flow is:
- Prepare input files: supports .pdf, .ofd, and image formats such as .jpg, .png, .bmp, .tif, .webp;
- Select a language code: for example, Simplified Chinese printed text uses 0, English uses 2, Japanese uses 11, and handwritten options map to other codes;
- Run batch recognition: files are sent to Xiangyun OCR, billed per call, and returned for further processing.

It fits engineering workflows that need multilingual documents normalized into text, especially scanned PDF/OFD files, receipts, and table screenshots.

Boundaries and Notes

The materials do not promise fully structured table output. If the goal is direct row and column extraction, treat OCR text as an intermediate result and add table parsing logic. Images should be clear, preferably larger than 500px in width and height; single files should stay under 10MB. The config file is stored as config.json in the skill directory. Credentials are sensitive and should not be committed to repositories or written to logs. Xiangyun OCR is billed per call, so validate cost and recognition quality on a small sample before running large batches.

Use Cases

  • Batch-recognize table text in scanned PDF or OFD files as input for field extraction
  • Convert printed Japanese, Korean, or Western-language documents from images into text for unified ingestion
  • After preprocessing invoice and form images, call OCR to obtain text for rule-based validation
  • Recognize text in multilingual scanned contracts before manual or programmatic review of key clauses

Best For

  • Document processing engineers who need to convert scanned PDF or OFD files into text for field extraction
  • Operations staff handling multilingual invoices and forms who need image text turned into validatable data
  • Backend developers integrating third-party OCR and managing credentials plus batch jobs
  • Compliance reviewers who need to recognize foreign-language contracts or tables before manual review