Web Content Extraction
Paste the following prompt into your AI chat to install this skill:
Follow https://skillhub.cn/install/skillhub.md to install @user_bb5d057c/web-content-extraction.
About this skill
Problem
When an engineer pastes a WeChat, Zhihu, API-doc, or skills.sh URL into an agent, the hard part is not opening the link; it is separating clean body content from navigation, lazy-loaded media, anti-scraping shells, and rendered UI. web-content-extraction treats the URL as input and produces Markdown text and media references that can be analyzed, summarized, or translated.
How it works
The skill selects tools by reliability and cost:
- Primary path: web_extract pulls article content into Markdown and fits ordinary web pages, some WeChat articles, and PDFs, handling up to 5 URLs per request.
- Discovery aid: web_search finds candidate pages by keyword or site: operator when no exact URL exists.
- Fallback path: when anti-bot controls, dynamic rendering, or login walls block extraction, browser_extract reads the page through browser_navigate and browser_snapshot.
For WeChat articles, the notes are explicit: repeated web_extract failures may be quota exhaustion rather than a site fault. A more reliable route runs scripts/fetch_weixin.sh to fetch HTML with curl, then parses the js_content region for body text, author, image data-src values, and media lists. If the user provides only keywords, the Sogou WeChat search flow resolves real URLs before full-text extraction.
Limits
This is a text-first extraction capability: body text and image URLs are usually recoverable, while video, audio, and image-only content may depend on JavaScript or lazy loading. Sogou search may rate-limit requests, and real URL resolution can fail occasionally. Content behind login and large PDFs that need splitting are outside scope. Extracted content is intended for the current conversation only and is not written to persistent memory.
Use Cases
- Hand a WeChat article link to the assistant to extract body text and author for a short competitive summary.
- Search Sogou WeChat by keyword, resolve the real URL, and fetch full text for hot-topic tracking.
- Read an API doc or skills.sh page into Markdown to compare interface fields and behavior.
- Fall back to browser snapshots when anti-bot controls or dynamic rendering block direct extraction.
Best For
- Operations analysts comparing competitor WeChat accounts, needing link text extracted to Markdown for summaries.
- Content editors tracking hot topics, needing keyword-based WeChat article search and full-text retrieval.
- Agent toolchain engineers needing stable webpage body extraction into analyzable text with fallback paths.
- Technical documentation researchers reading API docs or skill-market pages into structured key points.
Related Skills
Turns Taleb's antifragility, barbell strategy, and via negativa into executable prompts for risk analysis in career, investing, health, and other domains.
Keyword news search over ZAKER’s corpus with optional time filtering, returning up to 20 titles, authors, and summaries.
A generic knowledge review skill that generates role-based checklists, rrule-based reminders, and on-demand reviews without storing personal data.
A knowledge-management skill for layered retrieval across information sources.