DeepSeek Harness (DSH) plugin system allows AI assistant capabilities to be enhanced through extensions. In document processing scenarios, reading PDFs is a common requirement, but it often faces file size limits (e.g., 64KB) and the inability to directly read scanned or image-based pages. The dsh-pdf plugin uses pdfjs-dist to extract the complete Unicode text layer and automatically applies OCR to scanned/image pages, solving these pain points.

Plugin Introduction

The plugin is maintained by henryxiao709 and aims to allow AI assistants to read PDF files of any size. It extracts the complete Unicode text layer via pdfjs-dist, supporting Chinese, English, and other scripts; it also provides an automatic OCR fallback for scanned or image-based pages. The plugin is licensed under the MIT License.

Core Capabilities

  • Page-based reading: Supports specifying a reading range via the pages parameter, such as "1-3", "2", "1,3-5", or "all".
  • No file size limit: The default single-file limit is 200MB, breaking through the common 64KB limit.
  • Complete Unicode text layer: Uses pdfjs-dist to extract text, supporting multiple languages.
  • Automatic OCR fallback: In auto mode, if the number of characters in a page’s text layer is below the threshold, OCR is automatically applied.
  • Mode control: Provides three modes: text (text layer only, fast), ocr (forced OCR), and auto (automatic detection).
  • OCR engine selection: Supports Windows WinRT OCR (zero installation, requires specific language packs) or tesseract.js.

Installation and Enabling

Installation requires filesystem operations and a Node.js environment. Run the following commands to clone the repository and install dependencies, then inject it into a running DSH instance via command.

  1. Clone the repository:
    git clone https://github.com/henryxiao709/dsh-pdf.git
    cd dsh-pdf
  1. Install dependencies:
    npm install --ignore-scripts
  1. Inject into the running instance: In DSH’s Agent, execute the following command to inject the local path:
    dev_install_package <绝对路径到 dsh-pdf>
Alternatively, add the plugin to `dsh.profile.bundles` so it automatically loads when DSH starts.

Usage

The plugin registers the read_pdf tool, supporting parameters to control reading range and mode.

Tool signature:

read_pdf file_path=... [pages="1-3"] [mode=auto|text|ocr] [ocrEngine=auto|windows|tesseract] [maxCharsPerPage=20000]

Return example:

{
  "path": ".../test.pdf",
  "totalPages": 7,
  "mode": "auto",
  "pages": [
    { "number": 1, "text": "test… [OCR] test…", "source": "mixed", "chars": 163 },
    { "number": 2, "text": "1 test)…", "source": "text", "chars": 876 }
  ],
  "engines": ["windows: ok", "tesseract: unavailable (no traineddata)"],
  "warnings": []
}

Each page returns a source field identifying the data source: text (text layer only), ocr (OCR only), or mixed (mixed).

System Requirements and Notes

  • Runtime environment: Requires a running DeepSeek Harness instance and Node.js >= 20.
  • Windows OCR limitations: If using the Windows OCR engine, the system must have the zh-Hans-CN and en-US language packs installed.
  • Locked dependencies: The versions of @deepseek-ai/* dependencies are locked; no manual linking is required.
  • Configuration adjustments: Configuration items (such as maxFileBytes and maxCharsPerPage) can be adjusted in real time via Settings → Plugins → dsh-pdf.
  • License: MIT License.

dsh-pdf solves the file size limit and scan recognition issues for PDF reading in the DSH ecosystem, making it suitable for developers who need to process large numbers of documents or non-text PDFs locally. For more details, refer to the GitHub repository.