AI Agent Hub
Back to skills
💻

Browser Automation File Downloader Framework

Development Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @user_656c2e70/browser-file-downloader into my AI assistant according to https://skillhub.cn/install/skillhub.md.

About this skill

What problem it solves

Many institutional or government sites do not expose a direct file download link. The target content may depend on login state, cookies, sessions, or a print-to-PDF action; some tasks also require paging, search, filtering, and batch processing. In those cases, requests is often not enough.

How the skill works

The skill abstracts site-specific file download logic into a reusable browser-automation workflow. It first decides whether session persistence or a direct download link is required, then selects the most stable file path. The core strategies are ranked: direct download links, a Print page with page.pdf(), and a detail-page page.pdf() fallback. Implementation starts by diagnosing the DOM, then adapting URL patterns, waiting behavior, and window-opening behavior to the site type. For ASP.NET WebForms sites, the skill avoids page.goto() for detail pages to reduce the risk of losing __VIEWSTATE; opening the target URL with window.open(url, '_blank') in the same context is more stable. Waiting strategy prefers domcontentloaded and expect_download over networkidle, which can hang on ads or analytics. After download, the PDF should be validated so page guides or disclaimers are not mistaken for the target file.

Boundaries and cautions

If the file has a direct .pdf or .xlsx URL and no login is needed, requests is sufficient. For complex front-end rendering, pagination, or cross-entity lists, matching list-page text alone can produce false positives, so detail pages should be opened and verified. Print pages are usually cleaner than detail pages, but validation should ignore functional buttons such as Print or 列印 when checking for layout occlusion.

Use Cases

  • Download all historical disclosure PDFs for a given shareholder and stock from HKEX DION.
  • Batch-download filtered business reports from a logged-in ASP.NET WebForms portal after paging and search.
  • Extract PDF files from government pages that only offer print-to-PDF instead of direct links.
  • Validate downloaded PDFs and reject false positives from guides, disclaimers, or navigation occlusion.

Best For

  • Investment analysts who regularly download HKEX DION disclosure PDFs and verify their content.
  • Data engineers maintaining export scripts for logged-in internal portals with paging behavior.
  • Compliance staff organizing government form PDFs and verifying content completeness.
  • Automation engineers building reusable download scripts for newly adapted websites.