AI Agent Hub
Back to skills
Webpage to PDF and Markdown Converter icon

Webpage to PDF and Markdown Converter

Data Analysis Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @user_a0cd3220/url2pdf-mk according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

When capturing webpages or WeChat articles, teams often need a local copy that preserves text, images, and layout for offline reading, archiving, or editing. Manual copy-paste loses styling, plain downloads may only produce HTML or screenshots, and batch processing is more cumbersome.

How It Works

url2pdf-mk uses main.py as a unified entry point and routes based on input: one URL triggers single-page capture, while multiple URLs or an xlsx file trigger batch capture. If Chrome/Chromium/Edge is available, it launches a browser instance, uses CDP to retrieve the full DOM and computed styles, and calls the browser’s native PDF print pipeline to produce PDF + Markdown. Without a browser, it falls back to HTTP mode and produces Markdown.

  • Single-page mode: scrape.py captures a webpage while preserving images, layout, and styling.
  • Browser batch mode: batch_scrape.py starts one browser instance for multiple links in an xlsx file.
  • HTTP batch mode: batch_http.py avoids a browser and uses fewer resources, but cannot render JavaScript-driven content and does not produce PDF output.

It can also detect WeChat article titles and publish dates, then create date-named archive folders on the desktop for easier retrieval.

Boundaries and Notes

This skill is best suited for public webpages and WeChat articles that need offline copies, not for complex scraping or data pipelines. Videos and audio usually keep only cover images or links, WeChat mini programs are limited to static content, and HTTP mode may fail against anti-crawl protection. Default mode may reuse a real Chrome profile and access cookies or login sessions, so prefer --isolated for public content. Use login-dependent mode only when necessary, ideally in a trusted or sandboxed environment.

Use Cases

  • Save WeChat long articles to offline PDF and Markdown while keeping images and layout.
  • Read multiple article links from xlsx and batch-generate offline documents.
  • Extract static webpages without a browser and output Markdown only.
  • Detect WeChat article titles and publish dates, then archive them by date on the desktop.

Best For

  • Content editors who need to save WeChat articles or web resources offline.
  • Operations staff who batch-archive public account links while preserving layout.
  • Engineers who extract static webpages to Markdown in browserless environments.
  • Technical writers who convert webpages into searchable Markdown and PDF.