VA DocParse Document Parser
Paste the following prompt into your AI chat to install this skill:
Please install @user_15292d5a/yjkj-va-docparse according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem
Scanned PDFs and image-based documents often contain text, tables, and formulas that are hard to feed into downstream workflows. va-docparse targets this parsing scenario by calling a remote document parsing service through an MCP server. It converts PDF files, scanned PDFs, and common image formats into readable structured content, reducing the need for local OCR environments, custom parsing code, and multi-format recognition maintenance.
How It Works
The skill responds to intents such as OCR, document parsing, recognition, and reading, then calls the parse_document tool. The required input is an absolute file path, and the optional output_format can be markdown, json, or both. On success, the service returns extracted text content. On failure, it returns a coded error message, such as [F001] file not found, [F002] file unreadable, or [F003] unsupported format. For batch processing, each file is parsed independently; if the user requests merging, the results are concatenated in order with masked file names and failure codes preserved.
Boundaries
It supports PDF, scanned PDFs, PNG, JPG, and JPEG. It does not support audio, video, landscape photos, or person photos, and it does not directly support Word, Excel, or PPT; those files should be converted to PDF first. Results should be returned as the MCP service response without unsupported rewriting, completion, or expansion. When configuration is missing, network issues occur, or the service fails, the skill should report the error code and stop, without exposing MCP endpoints, keys, underlying commands, or other sensitive information.
Use Cases
- Reviewing scanned PDF contracts after converting their text and tables to Markdown.
- Extracting text, tables, and formulas from PNG/JPG screenshots for technical notes.
- Parsing multiple image-based records individually while preserving failure codes.
- Migrating an older Python parsing script to the MCP `parse_document` tool.
Best For
- Document engineers reviewing scanned contracts or reports who need editable text from PDFs.
- Application developers maintaining MCP workflows who need a reliable document parsing tool.
- Compliance staff handling screenshot or photo evidence who need to extract text and tables.
- Product operations staff entering data who need to batch convert image-based records to text.
Related Skills
Turns files, references, or chat context into styled single-page HTML reports with built-in templates and preset themes.
An engineering-oriented email automation solution for batch sending, Jinja2 templates, attachments, scheduled sending, inbox monitoring, and rule-based auto-reply.
An engineering workflow for creating, editing, reviewing, analyzing, and image-converting .docx files using pandoc, docx-js, OOXML, and LibreOffice.
Pre-submission scanner for Word/PDF blind-bid files that checks margins, fonts, page numbers, and metadata against tender requirements and flags residual bidder identities.