Web Assistant
Paste the following prompt into your AI chat to install this skill:
Follow https://skillhub.cn/install/skillhub.md and install @user_c9b69df4/web-assistant into your AI assistant.
About this skill
Problem
Many web pages are not static HTML: SPA routes are driven by JavaScript, list data arrives through XHR, and downloads may be protected by CDN referer checks. A plain snapshot or brittle selector strategy can fail quickly. Web Assistant reframes scraping as interface discovery and path validation, preferring direct API calls and URL-parameter-driven flows over fragile UI automation.
How It Works
The skill works in layers. It starts with navigate, evaluate, and snapshot to inspect page structure, network requests, and URL patterns. Extraction then follows API call > URL parameter > UI interaction. For authenticated sites, it checks session state, restores cookies, and persists state under auth/*.json. For protected downloads, it stays in the same origin, uses fetch to trigger the download, and stores files under ./Downloads/*. After research, stable API endpoints, selectors, and edge cases are recorded in references/sites/*/guide.md for reuse.
Boundaries
It is suited for structured site exploration, data extraction, file downloads, and site-guide maintenance. It requires browser access, file read/write, and capabilities such as run_code_unsafe. It is not a policy-free crawler: the material states it should not aggressively scrape, and endpoint versions, login state, and selectors still need evaluate checks per task.
Use Cases
- Extract paginated product detail data from an SPA list by locating the underlying XHR API instead of fragile selectors.
- Retrieve task data from a logged-in admin console after restoring the cookie session before running extraction.
- Save protected assets from a site that returns 403 to cross-origin fetches by triggering same-origin fetch and downloading to Downloads.
- Document multiple e-commerce detail-page structures into a site guide covering URL patterns, selectors, and API endpoints.
Best For
- Automation engineers who need to extract structured lists from multiple sites and want stable API endpoint or selector discovery.
- Data engineers extracting from logged-in backends who want to verify and restore session state before running jobs.
- Frontend or scraping engineers downloading protected assets who want same-origin fetch triggers to bypass 403 checks.
- Automation platform developers who want to persist site research into reusable site guides.
Related Skills
Fetches Baidu Hot Search Top 10 titles using web_fetch first, validates same-day data, and falls back to browser automation when stale.
Generates an evening A-share policy and trading opportunity daily report by collecting same-day index, policy, and capital data, then applying a fixed template to highlight beneficiary sectors, drivers, and price directions.
Maps natural-language TikTok requests to KeyAPI REST scenarios, validates endpoints against docs, and executes data queries and analysis.
Extract city-specified AI jobs from BOSS Zhipin, save CSV/table data, mark new postings, and summarize salary trends, application advice, and HTML reports.