Introduction¶
DeepSeek Harness (DSH) adopts a plugin-based architecture designed to encapsulate specific capabilities as pluggable components. When developing agents or processing documents, high-accuracy PDF parsing is a common requirement. Traditional parsing approaches often suffer from insufficient accuracy in layout analysis, formula recognition, or long-document processing, and may block the conversation. This plugin introduces the MinerU engine, providing DSH with native high-accuracy PDF reading capabilities.
What Is It?¶
dsh-pdf-mineru is a DSH plugin developed by Yurzi. As a provider-independent tool, it uses the MinerU engine to give agents document parsing capabilities. The plugin addresses difficulties AI faces in accurately extracting structured information such as formulas, tables, and layouts from complex PDFs (e.g., academic papers and research reports).
Core Features¶
The plugin offers the following core capabilities:
- Dual Deployment Support: Supports both MinerU’s official cloud (v4) and self-hosted private services (v2), keeping data within your internal network if required.
- High-Accuracy Parsing: Provides document layout analysis, formula and table extraction, and figure/text parsing.
- Asynchronous Processing and Caching: Natively supports background asynchronous tasks and intelligent caching, allowing synchronous responses or background processing; automatically deduplicates based on file fingerprints and configuration.
- Content Support: Supports double-column body text, tables of contents, LaTeX formulas, tables, references, and illustrations.
- Convenient Tooling: Provides a built-in visual configuration panel.
Installation and Enablement¶
In a DeepSeek Harness environment, run the following command to install the plugin:
dsh plugin --profile web add dsh-pdf-mineru
Typical Usage¶
Once configured, there is no need to memorize complex commands. You can simply state your requirements to the agent in natural language in the chat box. The plugin provides multiple interaction modes:
- Basic Parsing: Directly ask the agent to read a document and extract the main text, mathematical formulas, and tables, then organize them into Markdown format.
- Background Processing for Long Documents: For very long documents, use
async_parse_pdfto submit a task. The agent returns a task ID, parses the entire document into the local cache, and after completion returns a structured summary of the document (such as page count, outline, and number of figures/tables), without blocking the conversation. - Precise Slicing: Control the reading scope using parameters. For example, pass
pages: "1-5"together withfocus: "table"to read tables on specified pages, or passfocus: "toc"to view the table of contents. - Focused Reading and Retrieval: Use
query: "Fig. 7"for literal search; use the returnedblock_idto read a complete block in detail, or select content using a physical page number andfocus; for key formulas or figures, useview: "page"to return to the original page.
Use Cases and Notes¶
Before using this plugin, confirm that the following environment and configuration requirements are met:
- Environment Requirement: The runtime environment must satisfy Node.js
^22.19.0 || >=24.0.0. - DSH Version: Only supports RC (Release Candidate) versions of DeepSeek Harness and subsequent official release versions.
- Task Controller: The host must attach a task controller (such as
tool-jobs) to the session; otherwise, parsing is rejected at startup. - Error Handling: Errors such as corrupted files, encrypted files, parsing failures, and out-of-range page numbers are not retried by switching backends.
- Poppler Dependency: When both backends are unavailable, the plugin prompts you to install Poppler; original-page mode preferentially uses Poppler found in PATH; if missing or not executable, it automatically falls back to PDF.js + Node Canvas.
- Task Cancellation:
job_killonly cancels the current wait and does not terminate the shared parsing producer for other invocations.
Summary¶
By integrating the MinerU engine, dsh-pdf-mineru provides DeepSeek Harness with a PDF parsing solution that spans from cloud-based to local environments. Its asynchronous task processing, intelligent caching, and high-accuracy formula/table extraction capabilities can significantly improve the efficiency and accuracy of agents when processing documents.