AI Agent Hub
Back to skills
💻

Captcha Recognition Toolkit

Development Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please install @org-rn88lg3j/captcha according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

When automated scripts process web or app verification flows, they often need to convert captcha images into actionable outputs: text, gap coordinates, rotation angles, or click order. Captcha types and input formats vary, including image URLs and Base64 payloads. This skill groups those calls under a set of scripts.tools functions, allowing the model to route intent, extract parameters, and present results.

How It Works

  • Text and math recognition: find_chars returns characters and confidence; solve_math recognizes arithmetic and returns the computed result.
  • Slider recognition: slide_match_target_background, slide_compare, and related tools locate gaps or shadow-slider coordinates.
  • Rotation recognition: rotate_single and rotate_double return single-image or inner/outer relative rotation angles.
  • Ordered selection: detect_icon and detect_text detect icon or text coordinates according to a prompt image.
  • Calling convention: if XBY_APIKEY is missing, the skill must ask the user and save the key; incomplete parameters trigger follow-up questions; tools return result["raw"], result["success"], and result["message"] for presentation.

Boundaries

This skill fits developer use in authorized testing and compliant automation, especially when captcha images are already available but stable text or coordinates are needed. It does not bypass security controls and should not be used in violation of target-site terms. Without an API key, it must not search or fabricate results.

Use Cases

  • During authorized login tests, call find_chars to convert a text captcha image URL into text and confidence.
  • When a web slider challenge provides background and slider images, return the gap coordinate for dragging.
  • For rotation puzzles with single or inner/outer images, identify the 0-359 degree rotation angle.
  • For click-order challenges, detect icons or text from the prompt image and return ordered coordinates.

Best For

  • Test engineers automating login and checkout flows who need captcha images turned into text, coordinates, or angles.
  • Python developers building web scraping scripts who need to call recognition APIs for text, slider, rotation, and click-order captchas.
  • RPA engineers maintaining login pipelines who need gap backgrounds or prompt images converted into actionable coordinates.
  • QA tooling owners running compliant automation tests who need consistent routing for multiple captcha recognition calls and missing-parameter prompts.