DeepSeek Harness (DSH) adopts a plugin-based architecture and allows functionality to be extended as needed. When the main model (such as deepseek-official) has only text capabilities and cannot directly understand images, the dsh-umi-ocr-vision plugin acts as a visual bridge layer. It uses the locally running Umi-OCR tool to recognize image content and injects the recognized text as “untrusted observation data” into the request, enabling the text model to process textual information in images.

Core Features

  1. Automatic OCR bridging: When the main model has only text capabilities, the plugin automatically intercepts images in the chat, passes them to local Umi-OCR for text recognition, and then injects the results into the context.
  2. Multiple attachment support: Supports OCR processing of multiple attachments in the chat individually and submitting them in the same request.
  3. Data security: OCR results are marked as “untrusted observation data.” Images are not sent to any cloud service (HTTP mode also defaults to accessing only local 127.0.0.1).
  4. Toolset support: Provides a vision_* toolset (such as vision_ocr, vision_glance, vision_detect_text, etc.), allowing local OCR capabilities to be called during development or advanced configuration.

Installation

The plugin is provided as pure JS and requires no build.

  1. Install from a local source directory:
dsh plugin --profile web add /path/to/dsh-umi-ocr-vision
  1. Or package it into a tgz and then install:
dsh plugin --profile web add ./dsh-umi-ocr-vision-0.1.0.tgz

After installation, restart Harness for the plugin to take effect.

Prepare Umi-OCR

The plugin depends on the local Umi-OCR tool. Please download and start Umi-OCR first, and make sure its HTTP service is available.

  1. Download and start Umi-OCR.
  2. Enable the HTTP service in Umi-OCR, which by default listens on 127.0.0.1:1224.
  3. Optionally verify service availability: visit http://127.0.0.1:1224/api/ocr/get_options in a browser.

CLI Mode (Optional)

If you do not want to enable the HTTP service, you can configure the plugin to call the Umi-OCR command line directly.

In the llm-deepseek section of settings.yaml, set:

llm-deepseek:
  umiOcrMode: cli
  umiOcrCommand: "C:/Umi-OCR/Umi-OCR.exe"

The plugin first writes the image to a temporary directory, then executes the command line and reads the result.

Configuration

Configure OCR parameters in the llm-deepseek section of $DSH_HOME/settings.yaml:

llm-deepseek:
  umiOcrBaseURL: http://127.0.0.1:1224   # Umi-OCR HTTP 地址
  umiOcrMode: http                        # http 或 cli
  umiOcrCommand: ""                       # cli 模式时填写 Umi-OCR.exe 路径
  umiOcrTimeoutMs: 120000                 # 单张图片 OCR 超时(毫秒)
  maxImages: 8                            # 单次请求最多处理图片数
  cacheEntries: 64                        # OCR 结果缓存条数
  dataFormat: text                        # 桥接和工具统一使用 Umi-OCR dict 格式以获取文本坐标
  ocrLanguage: 简体中文                    # Umi-OCR 语言/模型库(Rapid 版用 "简体中文")
  ocrCls: false                           # 是否启用方向纠正
  ocrLimitSideLen: 960                    # 图像边长限制
  tbpuParser: multi_para                  # 布局解析方案
  enableVisionTools: true                 # 是否启用 vision_* 工具集
  artifactDir: .dsh-umi-vision/artifacts  # vision_* 工具产物输出目录
  longImageMaxHeight: 4096                # 长截图分块时单块最大高度
  longImageOverlap: 80                    # 长截图分块重叠像素

You can also modify the above configuration directly in Harness under Settings → Plugins → Plugin Configuration.

Toolset

After enabling enableVisionTools: true, the plugin registers the following tools with Harness:

  • vision_ocr: Performs OCR on an image and returns the full text and line-level coordinates.
  • vision_glance: Quickly inspects an image and returns its dimensions and visible text.
  • vision_detect_text: Lists all text blocks and their original coordinates.
  • vision_ground_text: Locates elements by text and returns x1,y1,x2,y2 coordinates.
  • vision_crop: Crops an image region and outputs a PNG.
  • vision_long_screenshot_ocr: Performs chunked OCR on long screenshots, merges Markdown, and saves the manifest.
  • vision_dominant_colors: Analyzes dominant colors and returns a HEX palette.
  • vision_pixel_diff: Compares two images and returns the difference percentage and difference regions.

sharp is an optional dependency. If it is not installed, OCR tools and automatic bridging remain available (automatic bridging degrades to direct whole-image OCR); however, image processing tools such as vision_crop, vision_dominant_colors, vision_pixel_diff, and vision_long_screenshot_ocr require sharp.

Use Cases and Notes

Use cases: This plugin is suitable for scenarios that require using local Umi-OCR to process screenshots, documents, CAPTCHAs, or interface text. If images require understanding spatial relationships or object semantics, consider using a true vision model plugin (such as dsh-vision).

Notes:
1. Permissions and security: The plugin runs with the permissions of the current DSH process. Before installing, check the source code and license (MIT).
2. Data trustworthiness: OCR results are marked as “untrusted observation data.” Prompts appearing in images do not gain system privileges.
3. Local execution: HTTP mode defaults to accessing only local 127.0.0.1 and does not involve cloud transmission.