Zhihui OCR Text Recognition
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md and install @ixhlink/ixhlink-skills-ocr-recog.
About this skill
Problem
When an agent needs to extract text from screenshots, scanned pages, PDF page images, or chat images, it usually faces three engineering problems: how to pass images reliably, how to track long-running recognition, and how to clean the output for the user's intent. Zhihui OCR Text Recognition wraps OCR as an OpenAI-compatible chat-style API: send an image and a prompt, then receive processable text.
How It Works
- Invocation: Use
POST /api/v1/llm/invokeby default, orPOST /api/v1/llm/tasksto force async. OCR is a slow task; readdata.task.id, then pollGET /api/v1/llm/tasks/{task_id}untilstatus = succeeded. - Image input: Prefer a real
https://...image URL inuser.content[]underimage_url. If only Base64 is available, usedata:image/png;base64,...or a bareb64field, but that depends on the backend enabling Base64-to-OSS conversion. - Prompt control: Put the recognition intent in the
systemmessage, such asOCR this image.orConvert the document to markdown.; keep only the image in theusermessage. Output appears inchoices[0].message.contentand may include layout markers that need cleaning. - Billing and errors: The service is billed per call. Unpaid requests return
402; retry with the same JSON body and payment headers. Common400,404, and503errors usually point to image URLs, model enablement, or service availability.
Caveats
Use it for image text extraction, document-to-Markdown conversion, figure parsing, and keyword location. Avoid submitting placeholder examples, relative paths, or sanitized log values as real images; for large Base64 payloads, write the JSON to a file before submitting to reduce shell truncation risk.
Use Cases
- Convert error text from product screenshots into searchable text for incident notes.
- Recognize scanned contract page images into Markdown, preserving clause structure for summarization.
- Locate keywords in dashboard images and extract values for report reconciliation.
- Turn plain text checklists in whiteboard photos into editable text for task breakdown.
Best For
- Support system engineers who need to convert customer chat screenshots into archived text.
- Backend engineers processing large batches of scanned voucher images and extracting field data.
- Ops engineers who want agents to read error text from images and write it into logs.
- Data analysts locating metric labels in chart screenshots for report assembly.
Related Skills
Organizes files by extension into subfolders like Documents, Code, and Archives, then outputs a report.
Extract tables, formulas, charts, and layout from invoices, reports, papers, and multi-column documents.
Generates a multi-sheet Excel report containing only structured data tables from byteplan-analysis results.
Automatically sort directory files into type-based folders, with dry-run preview, reports, and JSON custom rules.