AI Agent Hub
Back to skills
📁

PDF to Markdown

Office Efficiency Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md and install @org-rn88lg3j/pdf-md.

About this skill

Problem

PDF text, tables, and diagrams are often scattered across pages, and raw copying can produce misordered content, weak hierarchy, or lost visuals. PDF to Markdown targets office-work workflows by turning one PDF or a batch of same-folder PDFs into maintainable Markdown knowledge documents, reducing manual reformatting and excerpting, especially for tutorials, standards, textbooks, and internal docs.

How It Works & Limits

  • Structured extraction: It detects chapter hierarchy and outputs #, ##, and ### headings, lists, and tables, keeping heading levels, list indentation, and table rows/columns intact while preserving explicit source points without inventing content.
  • OCR and image handling: Scanned files require an XBY_APIKEY before OCR; flowcharts, architecture diagrams, table screenshots, and UI guides are retained as image references with nearby context notes.
  • Batch indexing: After processing a folder, it can generate a Wiki.md index with chapter links, relationships, and summary statistics; ambiguous hierarchy, terms, or naming is checked with the user.
  • Limitations: OCR output is machine-generated; poor-quality results may be marked [needs review] or [extraction failed], and it will not borrow content from other PDFs or existing notes to fill gaps.

Use Cases

  • Organize textbook chapters by converting PDF headings, lists, and tables into editable Markdown.
  • Process product manuals by recognizing scanned pages into Markdown while preserving figure context.
  • Archive internal standards by batch-processing same-folder PDFs and building a Wiki index.
  • Handle scanned files by recognizing page text after key setup and marking low-quality pages explicitly.

Best For

  • Graduate students organizing course PDFs into review notes
  • Engineers archiving vendor technical documents into a knowledge base
  • Administrative staff processing training materials and generating an index
  • Technical writers maintaining internal standard Markdown repositories