Introduction

The plugin mechanism of DeepSeek Harness (DSH) allows developers to extend the capabilities of agents through Node.js tools. When processing PDF documents, vision models are subject to token limits (approximately 384 tokens / 800×800 pixels equivalent). Rendering an entire page directly can lead to loss of details in small text, table structures, or vector graphics. To address this limitation, dsh-pdf-reader uses PyMuPDF to identify the content type of each page: text-heavy pages are processed through text extraction, while pages densely containing charts, tables, or formulas are rendered as high-DPI regions and then passed to the vision model for reading.

Plugin Overview

dsh-pdf-reader is a DeepSeek Harness plugin whose core function is content-aware PDF reading. It is built on PyMuPDF and can automatically apply different processing strategies based on page layout. If Python or required dependencies are missing, the tool returns clear warning messages (with specific installation commands), making it easier for agents or developers to repair the environment rather than causing the workflow to abort with a direct error.

Installation and Dependencies

To install the plugin, run the following command:

dsh plugin --profile web add dsh-pdf-reader

Before running the plugin tools, ensure the environment meets the following requirements:

  • Node.js: Version >= 20
  • Python: Version 3 or later
  • Python dependency: pymupdf (optional: pymupdf4llm)
  • DSH dependencies: @deepseek-ai/cordis ^4, @deepseek-ai/dsh-tools

Core Tools

The plugin provides three main tools for different stages of PDF processing.

pdf_scan

Used to obtain a content overview of a single page. It returns layout information for the page, including column structure, vector graphic regions, raster images, tables, text characters, presence of graphics, formula risk, and presence of a text layer. It is recommended to use this tool first to determine the page type before processing a specific page.

pdf_read_page –mode mixed

This is the core one-stop tool. When using --mode mixed, it performs the following actions:

  1. Low-resolution full-page preview: Returns a low-resolution full-page rendering (including formulas, table grid lines, illustration positions, and two-column layout order).
  2. Text extraction: Extracts the page’s text content.
  3. High-DPI chart/table rendering: Automatically detects chart and table regions and renders them as high-DPI PNG images.

All high-resolution images are cached in the .dsh-pdf-reader folder under the project root directory, and the tool returns only the image paths. This ensures that the context window is not filled with large amounts of image data.

pdf_render_region

Used for fine-grained rendering of specific regions. If the preview reveals that certain areas (such as formulas that were not automatically cropped, crowded table cells, or subplots) require more detail, this tool can be used. It accepts [x0, y0, x1, y1] coordinate parameters, renders the region at a DPI determined by the rendering budget, and returns the rendered image path.

Workflow Recommendations

The plugin implements a “preview -> content -> refinement” loop to save tokens and ensure that layout information is not lost:

  1. Preview: Call pdf_read_page --mode mixed to obtain the full-page preview image and text.
    • Use the preview image to inspect exactly what content is on the page and decide whether high-resolution rendering is needed.
  2. Content extraction: Call pdf_read_page --mode mixed again.
    • The tool automatically detects and renders all chart and table regions, writing them to the .dsh-pdf-reader cache. Pass these image paths to the vision model.
  3. Refinement: Call pdf_render_region as needed.
    • Zoom in on detail regions that were not automatically cropped in the preview and render them at a higher level of detail.

It is recommended to first use pdf_scan to get an overview of the entire document, planning which pages are suitable for text extraction and which pages require rendering.

Notes and Limitations

  • Table detection false positives: page.find_tables() may produce false positives on plot grids and charts. Therefore, tables are primarily handled through the rendering path to ensure structural accuracy, while Markdown tables are best-effort supplements.
  • Formula detection: Formula detection is based on heuristic rules (font + LaTeX producer). PyMuPDF parses inline mathematical formulas reasonably well, but complex structures such as fractions may need to be handled separately with pdf_render_region.
  • Memory usage: Large PDFs are parsed into memory.
  • Rendering strategy: The plugin favors “over-detection” rendering (that is, rendering grids or logos as if they were charts), because the rendering cost is relatively controllable and safe.

Conclusion

By separating text extraction from high-resolution rendering, dsh-pdf-reader addresses the token-limit issue faced by vision models when processing complex PDFs. It provides a toolchain from overview to fine-grained rendering, making it suitable for scenarios in DeepSeek Harness that require handling PDFs containing charts, formulas, and complex layouts. The plugin source code and directory can be accessed through the GitHub repository.