Windows Desktop Control
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @user_121462b9/desktop-control-win32.
About this skill
Problem
Many Windows apps, such as WeChat, QQ, and Electron apps, render UI through CEF/Chromium instead of standard Win32 child controls. As a result, pyautogui.click() can miss targets because of focus and hit-test differences. This skill targets the desktop loop of capturing a screenshot, analyzing it with a vision model, performing mouse and keyboard actions, then verifying the result.
How It Works And Limits
- Screenshot: captures a window or full-desktop image.
- Vision parsing: asks
GLM-4V-Flashfor percentage positions instead of exact pixel coordinates. - Control: uses
win32api.SetCursorPos,mouse_event, andkeybd_eventfor low-level clicks and keys. - Text input: short ASCII can use
pyautogui.write(); Chinese or emoji should be pasted via clipboard to reduce IME interference. - Verification: takes another screenshot and asks the vision model to confirm the action.
It is not intended for terminal, file I/O, or web scraping. Window handles may become invalid, Enter may send or insert a newline depending on app settings, and focus can be stolen, so delays, window re-enumeration, and final screenshot checks are important.
Use Cases
- Locate the input box in a Windows chat app, send a short message, and verify it with a follow-up screenshot.
- When CEF button clicks fail, use win32api with percentage-based coordinates from the vision model.
- Paste Chinese or emoji into desktop forms via clipboard to avoid IME interference.
- Complete click, type, and submit actions in QQ or Electron apps, then verify with screenshots.
Best For
- Automation engineers who need stable click, type, and message flows in WeChat, QQ, or other desktop apps.
- Test engineers validating UI actions in Electron or CEF applications with screenshot-based checks.
- Workflow developers integrating vision-based GUI understanding with low-level Windows input control.
- Multimodal agent engineers building screen-to-action loops for visible Windows windows.
Related Skills
Provides Claw with character-library selection, switching, saving, and global SOUL.md style sync for role-based conversation.
An AIONE Agentic AI Infrastructure SDK wrapper for building production AI agents with memory, skills, workflows, and hooks.
A Python/TypeScript SDK wrapper for the DeepSeek-Reasonix native AI coding agent, with prefix-cache support.
A browser automation tool for analysts, operators, and developers that locates elements, fills forms, extracts structured content, and supports no-code scheduling and export.