Scrapling Strategy Toolkit
Paste the following prompt into your AI chat to install this skill:
Please install @user_19a6d235/scrapling-fetcher-plus by following https://skillhub.cn/install/skillhub.md.
About this skill
Practical scraping workflow
When target pages mix static HTML, JS rendering, Cloudflare challenges, paginated lists, and auth state, one-off scripts quickly become brittle. This pack organizes Scrapling work into reusable strategies: pick Fetcher, DynamicFetcher, StealthyFetcher, or Spider by target type; extract fields with ::text, ::attr(name), :contains('...'), XPath, and regex fallback; then route output to CSV, JSON, Excel, WeCom, Lark, DingTalk, or IMA.
How it works and limits
The workflow is confirm-then-run: every task should state the URL, whether JS rendering is needed, whether Cloudflare is present, which fields to extract, and where the result should land. Static fetches suit public APIs, news, and docs; dynamic rendering suits SPAs and network-idle waits; StealthyFetcher targets stronger anti-bot checks; Spider fits product or search lists with concurrency, delay, and resume. Browser modes depend on Playwright and Chromium, platform routing may require credentials or external CLIs/Skills, and selectors can still break after front-end changes, so keep an OUTPUT_SCHEMA and preview step rather than hard-coding fragile scripts.
Use Cases
- Extract titles, article bodies, and links from public docs or news pages, then export CSV or Excel.
- Fetch SPA product prices or status after selector or network-idle waits, then write to Lark bitable.
- Fetch protected content behind Cloudflare or Turnstile challenges, then upload structured notes to IMA.
- Crawl paginated product or search lists with concurrency, delay, and resume, then summarize into WeCom sheets.
Best For
- Data engineers maintaining site monitors: reduce repeated selector rewrites with reusable templates.
- Operations staff archiving web content into enterprise knowledge bases: need structured extraction for IMA or WeCom docs.
- Product analysts tracking dynamic prices or inventory: need JS-render wait strategies and routing to Lark or WeCom.
- Researchers or scraping engineers handling Cloudflare-protected public content: need reusable StealthyFetcher parameter templates.
Related Skills
Fetches Baidu Hot Search Top 10 titles using web_fetch first, validates same-day data, and falls back to browser automation when stale.
Generates an evening A-share policy and trading opportunity daily report by collecting same-day index, policy, and capital data, then applying a fixed template to highlight beneficiary sectors, drivers, and price directions.
Maps natural-language TikTok requests to KeyAPI REST scenarios, validates endpoints against docs, and executes data queries and analysis.
Extract city-specified AI jobs from BOSS Zhipin, save CSV/table data, mark new postings, and summarize salary trends, application advice, and HTML reports.