Web Scraping Assistant
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md and install @user_b4da0893/web-scraping-helper-cn.
About this skill
Problem
Web scraping often stalls after the URL is collected: field definitions are unclear, selectors are unstable, pagination duplicates results, null or abnormal values are hard to audit, and handoff lacks acceptance criteria. This skill helps market research, operations, product, and data teams turn public page collection goals into executable and verifiable deliverables, not just general advice.
How It Works
It structures output into understanding confirmation, usable results, rationale, acceptance checklist, and next steps. Core capabilities include:
- Collection planning: start from
URL scopeand a field table for product pages, list pages, or detail pages. - Page parsing: define selection rules and fallbacks for titles, prices, parameters, and other fields.
- Pagination and deduplication: handle complete list capture with a pagination strategy and dedupe keys.
- Quality checks: inspect nulls, duplicates, anomalies, and missing fields.
- Change monitoring: track page changes by scrape frequency and changed fields.
The workflow usually begins by confirming that data is publicly accessible and reasonably usable, then limits domains, page types, frequency, and fields. It validates parsing rules on small samples, then adds deduplication, retries, rate limiting, and failure logs, and finally compares samples against the original pages with collection timestamps.
Boundaries
It does not bypass login, CAPTCHAs, paywalls, robots restrictions, or access controls, and it does not collect personal privacy data, sensitive identity information, or non-public data. When a request crosses boundaries, it should identify what cannot be done directly and provide compliant alternatives, required materials, or human handoff points. It fits scenarios with clear goals and source material where reviewers need checkable scripts and acceptance criteria, not requests to bypass protected platforms.
Use Cases
- Plan field tables, URL scope, and pagination dedupe rules for competitor product list research.
- Extract detail page titles, prices, and specs with selector rules and fallbacks.
- Audit collected data for nulls, duplicates, and abnormal prices, then produce an acceptance checklist.
- Monitor product page field changes by defining scrape frequency and watched changed fields.
Best For
- Market researchers planning competitor price research, field tables, and pagination dedupe rules.
- Operators maintaining product libraries who need detail-page title, price, and spec extraction with fallbacks.
- Product staff maintaining structured data who need null, duplicate, and anomaly checks.
- Data engineers running public collection tasks who need acceptance checklists and visible failures.
Related Skills
Fetches Baidu Hot Search Top 10 titles using web_fetch first, validates same-day data, and falls back to browser automation when stale.
Generates an evening A-share policy and trading opportunity daily report by collecting same-day index, policy, and capital data, then applying a fixed template to highlight beneficiary sectors, drivers, and price directions.
Maps natural-language TikTok requests to KeyAPI REST scenarios, validates endpoints against docs, and executes data queries and analysis.
Extract city-specified AI jobs from BOSS Zhipin, save CSV/table data, mark new postings, and summarize salary trends, application advice, and HTML reports.