PDF And Image Text Extraction
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @redfox-data/pdf-image-text-extractor-redfox.
About this skill
Problem
Text in documents is often trapped in screenshots, chat images, scanned pages, or complex PDFs. It may be hard to copy cleanly or feed into later search, cleanup, and knowledge-base workflows. Engineers need the headings, body text, notes, watermarks, and PDF paragraph hierarchy converted into editable Markdown, not transcribed by hand.
How It Works
The skill handles image and PDF text extraction. For images, it uses read_image and prompts the model to capture headings, body copy, notes, watermarks, and other text. For PDFs, it calls scripts/pdf_text_extractor.py to extract page text while preserving paragraph and heading structure as much as possible. It can return text directly or generate a .md file with source, extraction status, and content. For scanned pages, it notes that direct extraction may fail and suggests OCR.
Boundaries
Recognition quality depends on image clarity, fonts, and background. Tables, multi-column layouts, and scanned pages may return incomplete or empty text. Encrypted or password-protected PDFs are unsupported. Files under 50MB are recommended, and sensitive content is treated only within the current session. Complex tables or multi-column layouts should be reviewed manually.
Use Cases
- Convert order notes in chat screenshots into Markdown by extracting headings, body text, and annotations.
- Archive legacy PDF agreements by extracting page text while preserving heading levels into a searchable .md file.
- Check metric names and values in dashboard screenshots by recognizing the visible text and formatting it into a readable list.
- Inspect scanned PDFs to identify pages that cannot be extracted directly and get a note that OCR may be needed.
Best For
- Support engineers who need to turn user-provided screenshot notes into editable text for ticket archiving.
- Documentation engineers who need to convert legacy PDF guides into Markdown for search and citation.
- Data analysts who need to identify metric names, values, and annotations in dashboard screenshots.
- Legal assistants handling scanned contracts who need to know which PDF pages require OCR.
Related Skills
Organizes files by extension into subfolders like Documents, Code, and Archives, then outputs a report.
Extract tables, formulas, charts, and layout from invoices, reports, papers, and multi-column documents.
Generates a multi-sheet Excel report containing only structured data tables from byteplan-analysis results.
Automatically sort directory files into type-based folders, with dry-run preview, reports, and JSON custom rules.