AI Agent Hub
Back to skills
PDFlux Document to Markdown icon

PDFlux Document to Markdown

Knowledge Management Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Follow https://skillhub.cn/install/skillhub.md to install @user_15292d5a/yjkj-pdflux-pdf2markdown.

About this skill

Problem

Local documents entering LLM workflows often need reliable text order, table structure, and field extraction. pdflux-pdf2markdown narrows this to a single synchronous API call: send one local file and receive Markdown or JSON suitable for downstream processing.

How It Works

The skill expects direct invocation of scripts/upload_to_markdown.js rather than reimplementing the request flow. The script reads the Bearer key from PAODINGAI_API_KEY, posts the file to PDRouter's POST /openapi/{serviceCode}/file/markdown, prints Markdown when the response has a markdown field, and otherwise prints JSON. If output-markdown-path is provided, the same content is also written to that file. Progress and errors go to stderr, with a non-zero exit code on failure.

Boundaries

Charts and images are not included by default; enabling embedded images via PDFLUX_INCLUDE_IMAGES=true can increase token usage significantly. It works well as the first step for summarization, field extraction, table comparison, or rule-based validation. If the user asks for a direct conversion, return the Markdown; if only fields or tables are needed, filter from the parsed result instead of echoing the full output.

Use Cases

  • Parse a contract or report PDF into Markdown before summarizing it or extracting key fields.
  • When asked to convert a local document to Markdown, call the PDFlux synchronous API once and return the result.
  • Use the generated Markdown to read table fields for downstream rule-based validation or field extraction.
  • Prepare a single file as readable text at the start of a pipeline for table comparison or body extraction.

Best For

  • Document engineers who need contracts, reports, or manuals parsed into Markdown for summarization.
  • Data engineers who need to extract PDF table fields into automated processing scripts.
  • AI application developers building LLM workflows that require stable document text and table structure.
  • RAG system owners who need direct conversion output or only selected fields from documents.