PDF Scan to Markdown
Paste the following prompt into your AI chat to install this skill:
Please install @user_3afb9517/pdf-scan-to-md into your AI assistant following https://skillhub.cn/install/skillhub.md.
About this skill
Problem
Many technical docs, papers, and scanned files are not plain-text PDFs. Direct copying can lose table structure, math formulas, and multi-column layouts; image-heavy PDFs may not be extractable at all. pdf-scan-to-md targets these PDFs and converts text, formulas, and tables into editable Markdown.
How It Works
The skill relies on multi-engine conversion with fallback:
- Primary engine MinerU, suitable for scanned or image PDFs, math formulas to LaTeX, and table reconstruction; it can run locally or via cloud API.
- Secondary engine PaddleOCR, useful for Chinese/English text recognition, also available locally or online.
- Without tokens, it can use system OCR, tesseract, or prior local configurations as an offline fallback.
The typical flow is: prepare the PDF, optionally request a MinerU or PaddleOCR token, run the conversion, and let the script select an engine, switching if needed.
Boundaries
This skill is better suited to structured document conversion than simple text extraction. Cloud APIs usually produce better results, while local mode requires a model download and may be slower; MinerU local uses about 2GB and PaddleOCR local about 300MB. Files above limits may be split, and quotas depend on the provider.
Use Cases
- Convert scanned paper PDFs into editable Markdown with text, formulas, and tables.
- Extract clause text and rebuild tables from contracts under 100 pages for comparison.
- Convert multilingual scanned files into Markdown using local OCR or PaddleOCR offline.
- Process formula-heavy reports by converting math into LaTeX and checking structure.
Best For
- Engineers maintaining internal knowledge bases who need searchable Markdown from scans.
- Researchers organizing literature who need formulas, tables, and original structure preserved.
- Legal operations staff processing contract batches who need OCR extraction and table rebuilding.
- Developers running local AI pipelines who need offline OCR and engine fallback options.
Related Skills
Organizes files by extension into subfolders like Documents, Code, and Archives, then outputs a report.
Extract tables, formulas, charts, and layout from invoices, reports, papers, and multi-column documents.
Generates a multi-sheet Excel report containing only structured data tables from byteplan-analysis results.
Automatically sort directory files into type-based folders, with dry-run preview, reports, and JSON custom rules.