iFlytek BigModel OCR Image Understanding
Paste the following prompt into your AI chat to install this skill:
Please install @user_bd0773b4/iflytek-bigmodel-ocr according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem It Solves
When you already have a local image, screenshot, or table capture and need machine-readable content, use this skill to pass the image path directly instead of assembling a one-off OCR call. It fits office workflows where the result will be written into documents, spreadsheets, ticket fields, or downstream automation.
How It Works and Where It Binds
The skill first checks that the input image path exists, then resolves credentials and model settings. Configuration precedence is: CLI flags such as --api-key, --model-id, and --base-url first, then environment variables (default IFLYTEK_API_KEY), then the JSON file passed via --config. assets/config.template.json is only a template and should not be treated as real credentials. After execution, it returns a JSON result, and the content field can be explained when needed.
Use it for local images containing text, tables, or screenshot content. If a custom prompt is needed, pass it through arguments rather than editing files inside the skill package.
Use Cases
- Extract approval fields from a screenshot into JSON for ticketing
- Read table rows and cells from a local invoice image for reconciliation
- Pull error text from a product screenshot for troubleshooting records
- Pass a fixed model config via `--config` and run local image OCR
Best For
- Support-system developers who convert screenshot text into JSON
- Finance operations staff reconciling invoice table images
- QA engineers archiving product error screenshots into incident notes
- Automation script authors managing API keys via CLI flags
Related Skills
Turns files, references, or chat context into styled single-page HTML reports with built-in templates and preset themes.
An engineering-oriented email automation solution for batch sending, Jinja2 templates, attachments, scheduled sending, inbox monitoring, and rule-based auto-reply.
An engineering workflow for creating, editing, reviewing, analyzing, and image-converting .docx files using pandoc, docx-js, OOXML, and LibreOffice.
Pre-submission scanner for Word/PDF blind-bid files that checks margins, fonts, page numbers, and metadata against tender requirements and flags residual bidder identities.