AI Agent Hub
Back to skills
Windows Desktop Control icon

Windows Desktop Control

AI Agent Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_121462b9/desktop-control-win32.

About this skill

Problem

Many Windows apps, such as WeChat, QQ, and Electron apps, render UI through CEF/Chromium instead of standard Win32 child controls. As a result, pyautogui.click() can miss targets because of focus and hit-test differences. This skill targets the desktop loop of capturing a screenshot, analyzing it with a vision model, performing mouse and keyboard actions, then verifying the result.

How It Works And Limits

  • Screenshot: captures a window or full-desktop image.
  • Vision parsing: asks GLM-4V-Flash for percentage positions instead of exact pixel coordinates.
  • Control: uses win32api.SetCursorPos, mouse_event, and keybd_event for low-level clicks and keys.
  • Text input: short ASCII can use pyautogui.write(); Chinese or emoji should be pasted via clipboard to reduce IME interference.
  • Verification: takes another screenshot and asks the vision model to confirm the action.

It is not intended for terminal, file I/O, or web scraping. Window handles may become invalid, Enter may send or insert a newline depending on app settings, and focus can be stolen, so delays, window re-enumeration, and final screenshot checks are important.

Use Cases

  • Locate the input box in a Windows chat app, send a short message, and verify it with a follow-up screenshot.
  • When CEF button clicks fail, use win32api with percentage-based coordinates from the vision model.
  • Paste Chinese or emoji into desktop forms via clipboard to avoid IME interference.
  • Complete click, type, and submit actions in QQ or Electron apps, then verify with screenshots.

Best For

  • Automation engineers who need stable click, type, and message flows in WeChat, QQ, or other desktop apps.
  • Test engineers validating UI actions in Electron or CEF applications with screenshot-based checks.
  • Workflow developers integrating vision-based GUI understanding with low-level Windows input control.
  • Multimodal agent engineers building screen-to-action loops for visible Windows windows.