AI Agent Hub
Back to skills
arXiv Paper and Source Downloader icon

arXiv Paper and Source Downloader

Professional Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_c6886734/arxiv-paper-downloader-plus.

About this skill

Problem

Reading arXiv papers usually means handling several small tasks at once: extracting an ID from an abs/pdf URL, downloading the PDF or LaTeX source, and saving multiple papers with clean filenames. Manual copying and renaming becomes error-prone when titles contain :, /, or ?, or when titles are very long.

How It Works

The skill is built around scripts/download_arxiv.py, using only the Python standard library. It supports:

  • Single download: provide an arXiv ID, arxiv.org/abs/ID, or arxiv.org/pdf/ID; saves the PDF and renames it after the paper title
  • Keyword search: queries arXiv API results, then lets the user choose IDs to download
  • Source download: uses --source to fetch the .tar.gz LaTeX package
  • Batch download: processes multiple IDs in one command and reports success/failure counts

The workflow extracts IDs or URLs, determines the output directory, runs the script, and returns the saved path. Filename collisions get (1) or (2) appended; long titles are truncated at word boundaries; batch requests are spaced by 0.5s.

Notes

Not every paper has an available source, and arXiv may return 403. Metadata requests time out at 30s, downloads at 120s. If export.arxiv.org is blocked, retrying or changing the network may be necessary.

Use Cases

  • Download an arXiv PDF directly from an abs or pdf URL and keep the filename aligned with the paper title.
  • Search arXiv by keyword, compare metadata, and batch-save selected papers to a local directory.
  • Before reproducing an experiment, fetch the arXiv LaTeX source tarball to inspect code, figures, and assets.
  • Compile a literature set by downloading multiple arXiv papers in one pass while avoiding filename overwrites.

Best For

  • Researchers tracking new papers: save PDFs and sources by arXiv ID or URL for offline reading and reproduction.
  • Graduate students writing literature reviews: search by keyword, select candidates, then batch-download and tidy filenames.
  • Engineers reproducing models: fetch LaTeX source packages to extract code, configs, and figure assets.
  • Technical writers maintaining paper libraries: standardize filenames and avoid issues from special characters or long titles.