AI Agent Hub
Back to skills
Reference Harvester: Auto PDF & Citation Finder icon

Reference Harvester: Auto PDF & Citation Finder

Professional Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_82882ca7/reference-harvester.

About this skill

Problem

When managing papers or references, PDF reference sections often appear as numbered lists, author-first entries, or URL-only citations. Text extraction can also split https://, DOIs, and paths with spaces, producing unusable values like doi.org/ 10.1016/.... A simple regex pipeline may then yield truncated titles, missing authors, invalid BibTeX, and HTML error pages saved as PDFs.

How It Works

Reference Harvester turns the process into scripted steps: it extracts full text with pdfplumber, then parses References, 参考文献, or Bibliography sections using numbered and author-first reference patterns. It repairs broken URLs, removes page markers, headers, and footnote artifacts, and emits JSON with index, title, url, doi, arxiv_id, and raw. It then enriches each item via DOI or URL lookup to recover authors, journal, volume, and pages, and generates BibTeX entries as article for journal papers or unpublished for preprints. For downloads, it prioritizes arXiv, then attempts direct PDFs, PMC/EuropePMC, MDPI, Nature/Springer OA, PLoS, RSC, Wiley, and Elsevier endpoints, validating files by checking the %PDF- magic bytes.

Limitations

The skill is best for extracting references from a single paper PDF, not for resolving inline footnotes or paywalled items. Some publishers return 403 even for open-access papers, so manual download may still be needed. For nonstandard styles such as Vancouver or inline citations, the extracted text may need inspection and heuristic adjustments.

Use Cases

  • Extract references from a conference PDF, repair spaced DOIs, and generate a BibTeX file importable into Zotero.
  • Batch-download arXiv PDFs from parsed references and validate PDF headers so error pages are not saved as papers.
  • Try PMC, MDPI, and Springer sources for missing open-access PDFs, then export the unresolved reference list.
  • Create a landscape Word table of unresolved references with authors, titles, venues, and clickable download links.

Best For

  • Graduate students who need to batch-convert PDF references into BibTeX entries.
  • Research assistants maintaining a lab literature database and aligning OA PDFs with metadata.
  • Engineers verifying technical-report citations and confirming DOI or download-page links.
  • Postdocs preparing shared material packs and delivering unresolved reference lists.