AI Agent Hub
Back to skills
Chrome AI Action Browser Automation icon

Chrome AI Action Browser Automation

Development Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_e71390e8/chromeskill.

About this skill

Problem

When an AI needs to observe and operate a live web page, calling plain HTTP APIs is often not enough. Pages may render dynamically, require clicks, form input, reading interactive elements, capturing screenshots, or intercepting requests. chrome_skill turns these browser interactions into agent-callable actions, reducing the need to maintain ad-hoc scripts.

How It Works

It communicates through a local bridge to Chrome over CDP, supporting navigation, clicking, typing, screenshots, content extraction, Cookie management, LocalStorage, network interception, PDF export, and 60+ other actions. A typical workflow is: navigate to the page, waitForElement for key content, inspect the page with getText, getLinks, and getFormFields, then perform click, type, and pressKey, and finally verify with screenshot or evaluate. For unknown pages, start with getFocusableElements and getAccessibilityTree to discover interactive elements.

Scope and Notes

Use it for browser-based tasks, not for code generation, debugging, or work without page context. It expects local Chrome/Chromium and Node.js 18+. Non-ASCII URLs are usually encoded by the bridge; manually use encodeURIComponent only when terminal input causes garbling.

Use Cases

  • Debug a login form by opening it, typing credentials, clicking submit, and saving a screenshot.
  • Scrape a dynamic page by waiting for key elements, then extracting links, headings, and visible text.
  • Verify front-end changes by opening the target URL, checking console logs, network requests, and a viewport screenshot.
  • Analyze competitor pages by extracting form fields, focusable elements, and the accessibility tree for downstream modeling.

Best For

  • Front-end engineers running end-to-end tests who need reliable clicking, typing, navigation, and result screenshots.
  • Web-scraping engineers extracting text, links, form fields, and performance metrics from dynamic pages.
  • QA engineers maintaining automation who want reusable CDP actions for login, form submission, and Cookie checks.
  • Product engineers building browser extensions or local tools who need Chrome screenshot, PDF export, and network interception.