All-In-One OCR
Paste the following prompt into your AI chat to install this skill:
Please install @ixhlink/ixhlink-skills-quanneng-ocr-recog according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem
Images often contain text in screenshots, scans, tables, or figures. Manual transcription is inconsistent, and plain OCR may return text without layout or fail to support document-to-Markdown, figure parsing, and keyword localization. This skill targets a clear OCR pipeline: turn an image into editable text while preserving enough structure for downstream cleaning.
How It Works
All-In-One OCR exposes OCR as an asynchronous API. Submit POST /api/v1/llm/invoke or the task endpoint with OpenAI-compatible messages: put the scene prompt in system and only the image_url in user. Images can be provided as public https:// URLs, complete data: Base64 payloads, or raw b64 fields. The response returns a task_id; poll the task until succeeded, then read the output from choices[0].message.content.
- Document to Markdown: retain page structure.
- Plain OCR: extract text without layout assumptions.
- Figure parsing: handle diagrams and chart-like content.
- Keyword localization: locate specific terms in the image.
Output may include layout markers, so clean tags before displaying or indexing the text.
Boundaries
Base64 payloads must be real, decodable values, not example placeholders. If the model has not enabled backend conversion to OSS, Base64 submission can fail. Prefer public image URLs to avoid shell truncation of long Base64 strings. Usage is billed per call, and a 402 response requires retrying with the returned payment identifiers while keeping the original request body. It fits engineering use cases such as image text extraction, screenshot structuring, and document Markdown conversion; it is not a drop-in for offline deployment, high-volume instant batch processing, or environments without image access permissions.
Use Cases
- Convert supplier PDF screenshots into editable text for tables and body copy before downstream parsing.
- Extract order IDs, amounts, and complaint details from customer support screenshots into ticket fields.
- Pull axis labels, annotations, and captions from research figures into text for manual review.
- Transcribe campaign poster copy, prices, and dates into spreadsheet cells for review.
Best For
- Customer support leads who turn client screenshot complaints and order IDs into structured ticket text.
- Legal assistants who review scanned contracts and extract amounts, dates, and clause text.
- Technical writers who extract UI menu labels and prompts from screenshots into Markdown docs.
- Analysts who collect research material and transcribe posters, figures, and table screenshots.
Related Skills
Generate Markdown public opinion reports by calling an internal service with MIDU_API_KEY.
Extracts Google AI Mode answers, standard SERP, AI Overviews, and citations via Pangolin APIs, with multi-turn follow-ups and region support.
Provides break-even analysis frameworks and templates without code execution, outputting structured recommendations.
Generate web reports from existing analysis data with classic or PPT-style layouts, Chart.js charts, and keyboard/touch navigation.