AI Agent Hub
Back to skills
Smart Scraper Web icon

Smart Scraper Web

Data Analysis Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_15292d5a/yjkj-smart-scraper-web.

About this skill

Problem

Web fetch tools often return raw HTML, leaving agents to manually locate title, tables, prices, lists, and article text. Page structures vary, and repeated requests to the same page add latency.

How It Works

smart-scraper-web fetches pages using Node.js http/https and parses them with bounded regex patterns into structured output rather than full HTML. Core modes include --extract for title, headings, paragraphs, links, images, tables, lists, prices, and metadata; --table, --list, and --price for targeted extraction; --article for heading-plus-paragraph focus; --parse for already-fetched HTML; and --status for cache inspection. The cache is stored in memory/scraper-cache/cache.json, with a default 5-minute TTL and LRU eviction capped at 50 entries or 10MB. Security controls allow only public http/https URLs, block file://, data:, gopher://, localhost, private IPs, and cloud metadata endpoints, revalidate redirect targets, and limit redirects to 5.

Boundaries and Notes

It is not a DOM parser and does not execute JavaScript, so SPA content and frontend-rendered data are out of scope. Price detection is regex-based, recognizing $, €, £, and ¥ with thousands separators, but it does not infer meaning. Each page has a 15-second fetch timeout and a minimum 100ms interval between requests. After scraping sensitive pages, consider deleting cache.json.

Use Cases

  • Extract public pricing tables into rows and cells to build a competitor comparison sheet.
  • Pull product titles, prices, and image alt text from public pages to create basic comparison data.
  • Capture heading hierarchy and opening paragraphs from documentation pages to generate summary cards.
  • Parse already-fetched HTML into ordered and unordered lists to export menu structures.

Best For

  • Market research analysts: need to collect table and price fields from public pages into workable datasets.
  • Software engineers building agent toolchains: need to convert fetched pages into structured JSON for downstream models.
  • Content aggregation editors: need heading hierarchy and article previews from documentation pages.
  • Data engineers processing archived page snapshots: need offline parsing of lists and tables in existing HTML.