Webpage to PDF and Markdown Converter
Paste the following prompt into your AI chat to install this skill:
Please install @user_a0cd3220/url2pdf-mk according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem
When capturing webpages or WeChat articles, teams often need a local copy that preserves text, images, and layout for offline reading, archiving, or editing. Manual copy-paste loses styling, plain downloads may only produce HTML or screenshots, and batch processing is more cumbersome.
How It Works
url2pdf-mk uses main.py as a unified entry point and routes based on input: one URL triggers single-page capture, while multiple URLs or an xlsx file trigger batch capture. If Chrome/Chromium/Edge is available, it launches a browser instance, uses CDP to retrieve the full DOM and computed styles, and calls the browser’s native PDF print pipeline to produce PDF + Markdown. Without a browser, it falls back to HTTP mode and produces Markdown.
- Single-page mode:
scrape.pycaptures a webpage while preserving images, layout, and styling. - Browser batch mode:
batch_scrape.pystarts one browser instance for multiple links in anxlsxfile. - HTTP batch mode:
batch_http.pyavoids a browser and uses fewer resources, but cannot render JavaScript-driven content and does not produce PDF output.
It can also detect WeChat article titles and publish dates, then create date-named archive folders on the desktop for easier retrieval.
Boundaries and Notes
This skill is best suited for public webpages and WeChat articles that need offline copies, not for complex scraping or data pipelines. Videos and audio usually keep only cover images or links, WeChat mini programs are limited to static content, and HTTP mode may fail against anti-crawl protection. Default mode may reuse a real Chrome profile and access cookies or login sessions, so prefer --isolated for public content. Use login-dependent mode only when necessary, ideally in a trusted or sandboxed environment.
Use Cases
- Save WeChat long articles to offline PDF and Markdown while keeping images and layout.
- Read multiple article links from xlsx and batch-generate offline documents.
- Extract static webpages without a browser and output Markdown only.
- Detect WeChat article titles and publish dates, then archive them by date on the desktop.
Best For
- Content editors who need to save WeChat articles or web resources offline.
- Operations staff who batch-archive public account links while preserving layout.
- Engineers who extract static webpages to Markdown in browserless environments.
- Technical writers who convert webpages into searchable Markdown and PDF.
Related Skills
Fetches Baidu Hot Search Top 10 titles using web_fetch first, validates same-day data, and falls back to browser automation when stale.
Generates an evening A-share policy and trading opportunity daily report by collecting same-day index, policy, and capital data, then applying a fixed template to highlight beneficiary sectors, drivers, and price directions.
Maps natural-language TikTok requests to KeyAPI REST scenarios, validates endpoints against docs, and executes data queries and analysis.
Extract city-specified AI jobs from BOSS Zhipin, save CSV/table data, mark new postings, and summarize salary trends, application advice, and HTML reports.