Introduction

The plugin ecosystem of DeepSeek Harness (DSH) follows the philosophy of “everything is a plugin.” When using a pure-text model (such as DeepSeek), if the model needs to recognize an image, it typically tries to call the read_image tool. However, under pure-text routing, the backend does not support image input, so read_image will fail due to parameter mismatch.

dsh-ocr-vision solves this problem. It introduces a model-side ocr_image tool that uses a local RapidOCR engine to extract text from an image and return it as plain text, enabling pure-text models to “see” images.

Core Features

This plugin provides the following capabilities:
* Adds a new model-side ocr_image tool.
* Performs text extraction with local RapidOCR.
* Converts the extracted result into plain text for return.
* Supports reading images via ctx.fs.
* Supports region cropping.

Installation and Enablement

The plugin supports direct installation from GitHub as well as local build and installation.

Method 1: Direct Installation

Use the DSH plugin command to install directly from the GitHub repository:

dsh plugin --profile headless add github:QEDQCD/dsh-ocr-vision

After installation, dsh will automatically mount the ocr_image tool.

Method 2: Local Build Installation

  1. Clone the repository locally.
  2. Install dependencies and build:
   pnpm install && pnpm build
  1. Install to the specified profile:
   node scripts/install.mjs --profile headless

System Requirements

Before use, ensure that the host machine meets the following requirements:
* Operating system: Linux or macOS (Windows not tested).
* Python version: 3.7+.
* Python dependencies: rapidocr-onnxruntime, pillow, numpy.
* DeepSeek Harness: dsh CLI installed and a profile configured.

Verification

After installation, you can verify the setup in the following ways:
* Check the configuration: dsh --profile headless --dump-config | grep ocr-vision.
* Test the OCR script: python3 ocr.py <某图片路径>.

Usage

The model can directly call the ocr_image tool on the model side.

Basic call:

ocr_image(file_path: "截图.png")

Region cropping call:

ocr_image(file_path: "截图.png", region: "120,80,600,400")

The plain text returned by the tool has the following format:

<path>/workspace/截图.png</path>
<type>ocr</type>
<content>
报错码:ERR_DB_CONN_TIMEOUT
关键行:Connection refused -> db.internal:5432 / retry 3/3 failed
</content>

Configuration Items

The plugin supports adjusting behavior through cordis configuration:

Field Default Description
pythonBin python3 Path to the Python interpreter used to run the OCR script
ocrScript the ocr.py file included in the package Path to the OCR script
ocrCommand empty Full command prefix, used to override the default interpreter and script combination
scale 2 Image upscale factor (upscale by default to improve small-text recognition rate)
maxImageBytes 50 MiB Maximum image size in bytes when read via ctx.fs
timeoutMs 60000 Timeout for a single OCR subprocess in milliseconds

Limitations and Notes

  • Dependency requirements: The host must have Python and the RapidOCR dependencies installed; otherwise, ocr_image will report an error.
  • PDF not supported: Only raster image formats (such as PNG and JPEG) are supported; PDF input is not supported.
  • Does not intercept read_image: The plugin only adds the ocr_image tool. It does not intercept or rewrite read_image. If the model calls read_image under pure-text routing, it will still receive a tool rejection response.
  • Privacy and security: Images are read via ctx.fs and processed locally with OCR. Only the extracted text is returned to the model. Image bytes are not uploaded to the API.

Conclusion

dsh-ocr-vision provides localized and privacy-friendly image text recognition for pure-text models in DeepSeek Harness. It does not change how the model handles images; instead, it uses tool bridging to convert visual information into textual information.

  • Repository: https://github.com/QEDQCD/dsh-ocr-vision
  • Plugin directory: https://www.skillhub.cn/plugins/QEDQCD/dsh-ocr-vision