AI Agent Hub
Back to skills
Zhihui OCR Text Recognition icon

Zhihui OCR Text Recognition

Office Efficiency Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md and install @ixhlink/ixhlink-skills-ocr-recog.

About this skill

Problem

When an agent needs to extract text from screenshots, scanned pages, PDF page images, or chat images, it usually faces three engineering problems: how to pass images reliably, how to track long-running recognition, and how to clean the output for the user's intent. Zhihui OCR Text Recognition wraps OCR as an OpenAI-compatible chat-style API: send an image and a prompt, then receive processable text.

How It Works

  • Invocation: Use POST /api/v1/llm/invoke by default, or POST /api/v1/llm/tasks to force async. OCR is a slow task; read data.task.id, then poll GET /api/v1/llm/tasks/{task_id} until status = succeeded.
  • Image input: Prefer a real https://... image URL in user.content[] under image_url. If only Base64 is available, use data:image/png;base64,... or a bare b64 field, but that depends on the backend enabling Base64-to-OSS conversion.
  • Prompt control: Put the recognition intent in the system message, such as OCR this image. or Convert the document to markdown.; keep only the image in the user message. Output appears in choices[0].message.content and may include layout markers that need cleaning.
  • Billing and errors: The service is billed per call. Unpaid requests return 402; retry with the same JSON body and payment headers. Common 400, 404, and 503 errors usually point to image URLs, model enablement, or service availability.

Caveats

Use it for image text extraction, document-to-Markdown conversion, figure parsing, and keyword location. Avoid submitting placeholder examples, relative paths, or sanitized log values as real images; for large Base64 payloads, write the JSON to a file before submitting to reduce shell truncation risk.

Use Cases

  • Convert error text from product screenshots into searchable text for incident notes.
  • Recognize scanned contract page images into Markdown, preserving clause structure for summarization.
  • Locate keywords in dashboard images and extract values for report reconciliation.
  • Turn plain text checklists in whiteboard photos into editable text for task breakdown.

Best For

  • Support system engineers who need to convert customer chat screenshots into archived text.
  • Backend engineers processing large batches of scanned voucher images and extracting field data.
  • Ops engineers who want agents to read error text from images and write it into logs.
  • Data analysts locating metric labels in chart screenshots for report assembly.