Captcha Recognition Toolkit
Paste the following prompt into your AI chat to install this skill:
Please install @org-rn88lg3j/captcha according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem
When automated scripts process web or app verification flows, they often need to convert captcha images into actionable outputs: text, gap coordinates, rotation angles, or click order. Captcha types and input formats vary, including image URLs and Base64 payloads. This skill groups those calls under a set of scripts.tools functions, allowing the model to route intent, extract parameters, and present results.
How It Works
- Text and math recognition:
find_charsreturns characters and confidence;solve_mathrecognizes arithmetic and returns the computed result. - Slider recognition:
slide_match_target_background,slide_compare, and related tools locate gaps or shadow-slider coordinates. - Rotation recognition:
rotate_singleandrotate_doublereturn single-image or inner/outer relative rotation angles. - Ordered selection:
detect_iconanddetect_textdetect icon or text coordinates according to a prompt image. - Calling convention: if
XBY_APIKEYis missing, the skill must ask the user and save the key; incomplete parameters trigger follow-up questions; tools returnresult["raw"],result["success"], andresult["message"]for presentation.
Boundaries
This skill fits developer use in authorized testing and compliant automation, especially when captcha images are already available but stable text or coordinates are needed. It does not bypass security controls and should not be used in violation of target-site terms. Without an API key, it must not search or fabricate results.
Use Cases
- During authorized login tests, call find_chars to convert a text captcha image URL into text and confidence.
- When a web slider challenge provides background and slider images, return the gap coordinate for dragging.
- For rotation puzzles with single or inner/outer images, identify the 0-359 degree rotation angle.
- For click-order challenges, detect icons or text from the prompt image and return ordered coordinates.
Best For
- Test engineers automating login and checkout flows who need captcha images turned into text, coordinates, or angles.
- Python developers building web scraping scripts who need to call recognition APIs for text, slider, rotation, and click-order captchas.
- RPA engineers maintaining login pipelines who need gap backgrounds or prompt images converted into actionable coordinates.
- QA tooling owners running compliant automation tests who need consistent routing for multiple captcha recognition calls and missing-parameter prompts.
Related Skills
Analyzes code to extract control and data flow, then outputs Markdown with Mermaid source and high-resolution PNG diagrams.
For development and programming scenarios around VSCode and TypeScript IDE.
A TypeScript-oriented Windmill Wrap development reference.
A Python-based Selenium wrapper for engineering browser automation workflows.