Introduction¶
DeepSeek Harness (DSH) extends its capabilities through the plugin mechanism. When developing DSH agents, processing PDF documents is a common requirement. However, standard PDF rendering engines (such as pdfjs) often fail to handle PDFs with non-embedded CJK fonts. This is common in schematics exported from JLCPCB or EasyEDA, causing missing text or garbled characters.
zhtx2024/dsh-pdf addresses this issue by providing DSH with a set of PDF parsing tools. Through a hybrid engine design, it resolves difficulties in rendering and extracting PDFs with non-embedded Chinese fonts.
Core Features¶
The plugin provides three native tools for DSH agents and includes built-in dual-engine rendering logic:
-
pdf_info
Retrieves metadata for a PDF document, including total page count, page size in points (pt), document generation information, and a font list. The font list indicates the embedded status (embedded) of each font. -
pdf_extract_text
Extracts text content from the PDF. Supports specifying page numbers or extraction position coordinates (x/y coordinates and font size for each line) via parameters. -
pdf_render_page
Renders the specified page to a PNG image. Supports custom scale ratio (scale) and output directory (outDir).
Installation and Activation¶
Install the plugin using DSH’s command-line tool. After installation, restart DSH or start a new session for the tools to take effect.
dsh plugin --profile web add @zhtx2026/dsh-pdf
Typical Usage¶
After installation, you can directly call the tools mentioned above in a DSH session.
- Insight document information
pdf_info("D:/Downloads/SCH_SA-V11A.pdf")
- Extract text from a specified page
pdf_extract_text("D:/Downloads/SCH_SA-V11A.pdf", { page: 1 })
- Render a high-resolution page
pdf_render_page("D:/Downloads/SCH_SA-V11A.pdf", { page: 1, scale: 2 })
Engines and Principles¶
This plugin adopts a hybrid engine architecture and automatically switches rendering modes based on the PDF font situation:
- pdfjs engine: When all fonts in the PDF are embedded, standard pdfjs is used for text extraction and rendering. This is the default method for processing most standard PDFs.
- sysfont engine: For PDFs with non-embedded CJK fonts (such as schematics exported from JLCPCB), the plugin automatically enables the built-in decoder. This decoder parses PDF content streams (such as
BT/ET,Tf,Tmoperators), mapsUniGB-UCS2-Hor UTF-16BE encoded data to Unicode, and calls system fonts (such as SimSun and SimHei) to draw the text. This avoids thetranslateFont failederror thrown by pdfjs when it encounters non-embedded fonts.
Notes¶
The following limitations should be noted:
- Scanned Document Limitation: If the PDF is purely image-based or a scan (with no text layer),
pdf_extract_textwill return an empty result, butpdf_render_pageremains usable. - Font Assumption: The
sysfontrenderer assumes thatType0CID codes are equal to Unicode code points (forUniGB-UCS2-Hfonts). Other special CID fonts fall back to the pdfjs engine. - Memory Usage: Large files are fully parsed into memory.
- Session Restart: After installation, restarting DSH or opening a new session is required for the hot-plug mounting of the plugin to take effect.
Conclusion¶
zhtx2024/dsh-pdf closes the gap in DeepSeek Harness regarding text rendering for PDFs exported by Chinese EDA tools through its dual-engine strategy. To view the source code or catalog information, visit GitHub or Community Catalog.