Browser Automation File Downloader Framework
Paste the following prompt into your AI chat to install this skill:
Please install @user_656c2e70/browser-file-downloader into my AI assistant according to https://skillhub.cn/install/skillhub.md.
About this skill
What problem it solves
Many institutional or government sites do not expose a direct file download link. The target content may depend on login state, cookies, sessions, or a print-to-PDF action; some tasks also require paging, search, filtering, and batch processing. In those cases, requests is often not enough.
How the skill works
The skill abstracts site-specific file download logic into a reusable browser-automation workflow. It first decides whether session persistence or a direct download link is required, then selects the most stable file path. The core strategies are ranked: direct download links, a Print page with page.pdf(), and a detail-page page.pdf() fallback. Implementation starts by diagnosing the DOM, then adapting URL patterns, waiting behavior, and window-opening behavior to the site type. For ASP.NET WebForms sites, the skill avoids page.goto() for detail pages to reduce the risk of losing __VIEWSTATE; opening the target URL with window.open(url, '_blank') in the same context is more stable. Waiting strategy prefers domcontentloaded and expect_download over networkidle, which can hang on ads or analytics. After download, the PDF should be validated so page guides or disclaimers are not mistaken for the target file.
Boundaries and cautions
If the file has a direct .pdf or .xlsx URL and no login is needed, requests is sufficient. For complex front-end rendering, pagination, or cross-entity lists, matching list-page text alone can produce false positives, so detail pages should be opened and verified. Print pages are usually cleaner than detail pages, but validation should ignore functional buttons such as Print or 列印 when checking for layout occlusion.
Use Cases
- Download all historical disclosure PDFs for a given shareholder and stock from HKEX DION.
- Batch-download filtered business reports from a logged-in ASP.NET WebForms portal after paging and search.
- Extract PDF files from government pages that only offer print-to-PDF instead of direct links.
- Validate downloaded PDFs and reject false positives from guides, disclaimers, or navigation occlusion.
Best For
- Investment analysts who regularly download HKEX DION disclosure PDFs and verify their content.
- Data engineers maintaining export scripts for logged-in internal portals with paging behavior.
- Compliance staff organizing government form PDFs and verifying content completeness.
- Automation engineers building reusable download scripts for newly adapted websites.
Related Skills
A one-shot coding agent built on Claude Code CLI that runs non-interactively, supports a specified workdir, and can be monitored in the foreground or background.
Preview and confirm file sorting by extension, with recursive cleanup, ignore rules, and transactional rollback.
An engineering assistant for static HTML/CSS/JS pages, design-token extraction, IE8-compatible review, and structured delivery.
An engineering workflow for requirement analysis, scenario modeling, risk planning, quality gates, testing, and knowledge capture, with lightweight, standard, and full modes.