Web Scraping and Understanding
Paste the following prompt into your AI chat to install this skill:
Please install @user_f12a44b7/web-scraper-cascade according to https://skillhub.cn/install/skillhub.md.
About this skill
The Problem
Web scraping often fails because pages vary in structure, rely on JavaScript rendering, use anti-bot controls, or sit behind paywalls. Sending raw HTML to an LLM is also noisy and expensive. This skill treats extraction as a staged engineering task, which is useful when you need reliable JSON from articles, metadata, or public entity data.
How It Works
- Planning first: identify target URLs, content scope, crawl size, and output format before generating code.
- Five-stage pipeline: detect article-like content, fetch over HTTP, parse HTML, render with Playwright when needed, and optionally extract entities with an LLM.
- Config-driven selectors: maintain CSS selectors and pipeline settings in
YAML, while generatedPythonscripts emitJSONoutput. - Credential safety: use
os.environ.get()in generated scripts and avoid hard-coded keys, authentication bypasses, or hard paywall circumvention.
Boundaries
It is best suited to public articles, documentation pages, metadata extraction, and entity extraction. It does not promise to bypass login pages, paid content, or aggressive anti-bot systems. Stage 5 requires OPENROUTER_API_KEY; earlier stages generally need no credentials. Frequent site redesigns still require selector and output-quality monitoring.
Use Cases
- Batch-extract titles, publish dates, and article text from public pages into unified JSON.
- Check whether a site needs JavaScript rendering, then generate rerunnable Playwright scripts.
- Maintain CSS selectors in YAML and periodically extract metadata from public documentation pages.
- Extract entities from public pages in scripts, returning JSON with an error field on failure.
Best For
- Data engineers needing structured JSON from public web articles
- Scraping engineers maintaining selector configs and rerunnable crawl scripts
- ML engineers extracting entities with graceful LLM failure outputs
- Analysts collecting metadata from public documentation sites
Related Skills
Generate Markdown public opinion reports by calling an internal service with MIDU_API_KEY.
Extracts Google AI Mode answers, standard SERP, AI Overviews, and citations via Pangolin APIs, with multi-turn follow-ups and region support.
Provides break-even analysis frameworks and templates without code execution, outputting structured recommendations.
Generate web reports from existing analysis data with classic or PPT-style layouts, Chart.js charts, and keyboard/touch navigation.