PDFlux Document to Markdown
Paste the following prompt into your AI chat to install this skill:
Follow https://skillhub.cn/install/skillhub.md to install @user_15292d5a/yjkj-pdflux-pdf2markdown.
About this skill
Problem
Local documents entering LLM workflows often need reliable text order, table structure, and field extraction. pdflux-pdf2markdown narrows this to a single synchronous API call: send one local file and receive Markdown or JSON suitable for downstream processing.
How It Works
The skill expects direct invocation of scripts/upload_to_markdown.js rather than reimplementing the request flow. The script reads the Bearer key from PAODINGAI_API_KEY, posts the file to PDRouter's POST /openapi/{serviceCode}/file/markdown, prints Markdown when the response has a markdown field, and otherwise prints JSON. If output-markdown-path is provided, the same content is also written to that file. Progress and errors go to stderr, with a non-zero exit code on failure.
Boundaries
Charts and images are not included by default; enabling embedded images via PDFLUX_INCLUDE_IMAGES=true can increase token usage significantly. It works well as the first step for summarization, field extraction, table comparison, or rule-based validation. If the user asks for a direct conversion, return the Markdown; if only fields or tables are needed, filter from the parsed result instead of echoing the full output.
Use Cases
- Parse a contract or report PDF into Markdown before summarizing it or extracting key fields.
- When asked to convert a local document to Markdown, call the PDFlux synchronous API once and return the result.
- Use the generated Markdown to read table fields for downstream rule-based validation or field extraction.
- Prepare a single file as readable text at the start of a pipeline for table comparison or body extraction.
Best For
- Document engineers who need contracts, reports, or manuals parsed into Markdown for summarization.
- Data engineers who need to extract PDF table fields into automated processing scripts.
- AI application developers building LLM workflows that require stable document text and table structure.
- RAG system owners who need direct conversion output or only selected fields from documents.
Related Skills
Search Huawei Cloud official docs and product pages to find ECS, OBS, RDS, CCE product specs, parameters, documentation, and API references without login.
OCR-based recognition for movie, train, flight, and event tickets in images or PDFs, extracting key fields into Markdown or JSON reports.
Turns notes, research, and meeting summaries into actionable next moves, plans, decisions, experiments, and decision-changing gaps.
A local wiki knowledge base manager that compiles raw documents into sourced, indexed Markdown pages with wikilinks, query support, and health checks.