Introduction

DeepSeek Harness (DSH) extends its capabilities through the plugin mechanism. When developing DSH agents, processing PDF documents is a common requirement. However, standard PDF rendering engines (such as pdfjs) often fail to handle PDFs with non-embedded CJK fonts. This is common in schematics exported from JLCPCB or EasyEDA, causing missing text or garbled characters.

zhtx2024/dsh-pdf addresses this issue by providing DSH with a set of PDF parsing tools. Through a hybrid engine design, it resolves difficulties in rendering and extracting PDFs with non-embedded Chinese fonts.

Core Features

The plugin provides three native tools for DSH agents and includes built-in dual-engine rendering logic:

  1. pdf_info
    Retrieves metadata for a PDF document, including total page count, page size in points (pt), document generation information, and a font list. The font list indicates the embedded status (embedded) of each font.

  2. pdf_extract_text
    Extracts text content from the PDF. Supports specifying page numbers or extraction position coordinates (x/y coordinates and font size for each line) via parameters.

  3. pdf_render_page
    Renders the specified page to a PNG image. Supports custom scale ratio (scale) and output directory (outDir).

Installation and Activation

Install the plugin using DSH’s command-line tool. After installation, restart DSH or start a new session for the tools to take effect.

dsh plugin --profile web add @zhtx2026/dsh-pdf

Typical Usage

After installation, you can directly call the tools mentioned above in a DSH session.

  1. Insight document information
    pdf_info("D:/Downloads/SCH_SA-V11A.pdf")
  1. Extract text from a specified page
    pdf_extract_text("D:/Downloads/SCH_SA-V11A.pdf", { page: 1 })
  1. Render a high-resolution page
    pdf_render_page("D:/Downloads/SCH_SA-V11A.pdf", { page: 1, scale: 2 })

Engines and Principles

This plugin adopts a hybrid engine architecture and automatically switches rendering modes based on the PDF font situation:

  • pdfjs engine: When all fonts in the PDF are embedded, standard pdfjs is used for text extraction and rendering. This is the default method for processing most standard PDFs.
  • sysfont engine: For PDFs with non-embedded CJK fonts (such as schematics exported from JLCPCB), the plugin automatically enables the built-in decoder. This decoder parses PDF content streams (such as BT/ET, Tf, Tm operators), maps UniGB-UCS2-H or UTF-16BE encoded data to Unicode, and calls system fonts (such as SimSun and SimHei) to draw the text. This avoids the translateFont failed error thrown by pdfjs when it encounters non-embedded fonts.

Notes

The following limitations should be noted:

  1. Scanned Document Limitation: If the PDF is purely image-based or a scan (with no text layer), pdf_extract_text will return an empty result, but pdf_render_page remains usable.
  2. Font Assumption: The sysfont renderer assumes that Type0 CID codes are equal to Unicode code points (for UniGB-UCS2-H fonts). Other special CID fonts fall back to the pdfjs engine.
  3. Memory Usage: Large files are fully parsed into memory.
  4. Session Restart: After installation, restarting DSH or opening a new session is required for the hot-plug mounting of the plugin to take effect.

Conclusion

zhtx2024/dsh-pdf closes the gap in DeepSeek Harness regarding text rendering for PDFs exported by Chinese EDA tools through its dual-engine strategy. To view the source code or catalog information, visit GitHub or Community Catalog.