Introduction¶
The plugin ecosystem of DeepSeek Harness (DSH) follows the philosophy of “everything is a plugin.” When using a pure-text model (such as DeepSeek), if the model needs to recognize an image, it typically tries to call the read_image tool. However, under pure-text routing, the backend does not support image input, so read_image will fail due to parameter mismatch.
dsh-ocr-vision solves this problem. It introduces a model-side ocr_image tool that uses a local RapidOCR engine to extract text from an image and return it as plain text, enabling pure-text models to “see” images.
Core Features¶
This plugin provides the following capabilities:
* Adds a new model-side ocr_image tool.
* Performs text extraction with local RapidOCR.
* Converts the extracted result into plain text for return.
* Supports reading images via ctx.fs.
* Supports region cropping.
Installation and Enablement¶
The plugin supports direct installation from GitHub as well as local build and installation.
Method 1: Direct Installation¶
Use the DSH plugin command to install directly from the GitHub repository:
dsh plugin --profile headless add github:QEDQCD/dsh-ocr-vision
After installation, dsh will automatically mount the ocr_image tool.
Method 2: Local Build Installation¶
- Clone the repository locally.
- Install dependencies and build:
pnpm install && pnpm build
- Install to the specified profile:
node scripts/install.mjs --profile headless
System Requirements¶
Before use, ensure that the host machine meets the following requirements:
* Operating system: Linux or macOS (Windows not tested).
* Python version: 3.7+.
* Python dependencies: rapidocr-onnxruntime, pillow, numpy.
* DeepSeek Harness: dsh CLI installed and a profile configured.
Verification¶
After installation, you can verify the setup in the following ways:
* Check the configuration: dsh --profile headless --dump-config | grep ocr-vision.
* Test the OCR script: python3 ocr.py <某图片路径>.
Usage¶
The model can directly call the ocr_image tool on the model side.
Basic call:
ocr_image(file_path: "截图.png")
Region cropping call:
ocr_image(file_path: "截图.png", region: "120,80,600,400")
The plain text returned by the tool has the following format:
<path>/workspace/截图.png</path>
<type>ocr</type>
<content>
报错码:ERR_DB_CONN_TIMEOUT
关键行:Connection refused -> db.internal:5432 / retry 3/3 failed
</content>
Configuration Items¶
The plugin supports adjusting behavior through cordis configuration:
| Field | Default | Description |
|---|---|---|
pythonBin |
python3 |
Path to the Python interpreter used to run the OCR script |
ocrScript |
the ocr.py file included in the package |
Path to the OCR script |
ocrCommand |
empty | Full command prefix, used to override the default interpreter and script combination |
scale |
2 |
Image upscale factor (upscale by default to improve small-text recognition rate) |
maxImageBytes |
50 MiB |
Maximum image size in bytes when read via ctx.fs |
timeoutMs |
60000 |
Timeout for a single OCR subprocess in milliseconds |
Limitations and Notes¶
- Dependency requirements: The host must have Python and the RapidOCR dependencies installed; otherwise,
ocr_imagewill report an error. - PDF not supported: Only raster image formats (such as PNG and JPEG) are supported; PDF input is not supported.
- Does not intercept
read_image: The plugin only adds theocr_imagetool. It does not intercept or rewriteread_image. If the model callsread_imageunder pure-text routing, it will still receive a tool rejection response. - Privacy and security: Images are read via
ctx.fsand processed locally with OCR. Only the extracted text is returned to the model. Image bytes are not uploaded to the API.
Conclusion¶
dsh-ocr-vision provides localized and privacy-friendly image text recognition for pure-text models in DeepSeek Harness. It does not change how the model handles images; instead, it uses tool bridging to convert visual information into textual information.
- Repository: https://github.com/QEDQCD/dsh-ocr-vision
- Plugin directory: https://www.skillhub.cn/plugins/QEDQCD/dsh-ocr-vision