Preface

DeepSeek Harness (DSH) uses a plugin-based architecture to enable modular system capability extensions. Pure text models (such as DeepSeek) cannot directly process image inputs. dsh-omni-vision is a plugin maintained by Renji004. It provides pure text models with “local eyes,” allowing Agents to indirectly perceive visual content through rendering, OCR, and pixel analysis. The entire process runs locally and does not rely on cloud-based vision models.

Core Features

The plugin provides four core tools, corresponding to four actions: draw, paste, read, and see.

eyes_render

Renders a canvas in the Web GUI, supporting text, lines, rectangles, circles, and ellipses. The rendering result generates a PNG file, saved to the renders/ directory in the workspace (automatically deleted after being read, i.e., “burn after reading”) and to the pics/ directory (permanently retained).

eyes_paste

Allows users to paste an image from the clipboard or directly drag and drop an image into the browser interface. The image is captured and saved as a PNG to the renders/ directory for subsequent processing, without generating a permanent copy.

eyes_ocr

Uses the built-in Windows OCR engine (Windows.Media.Ocr) to recognize text in a PNG image and return it as text. This is the primary way for a pure text model to “see text.”

eyes_analyze

Decodes a PNG into pixel data and supports configuring an N×N sampling grid. It returns a color grid, a dominant color histogram, and a content bounding box, allowing the model to “see” shapes and color distributions.

Installation and Activation

Run the following command in the DSH workspace directory to install the plugin:

cd <your_workspace>
dsh plugin --profile web add link:<your_workspace>\dsh-omni-vision

After installation, the Web service must be restarted for the plugin to take effect.

Typical Usage

The following are typical application scenarios for this plugin in the DSH workflow:

  1. OCR Recognition: Use eyes_paste to let the user paste a screenshot, then use eyes_ocr to read the text inside it and paraphrase it.
  2. Drawing Recognition: Use eyes_render to draw an 800x400 canvas, write the text “Hello World 123” on it, set ocr:true, and then paraphrase the text seen.
  3. Flowchart Analysis: Use eyes_render to draw a Mermaid flowchart (such as graph TD...), set ocr:true, and read the node text.
  4. Text Reiteration: Use eyes_ocr to read <workspace>\renders\eyes-xxx.png, and reiterate the recognized text line by line.
  5. Pixel Analysis: Use eyes_analyze to analyze <workspace>\renders\eyes-xxx.png, set the grid to 12, and describe the layout and dominant colors of the image.

Limitations and Notes

The following limitations should be noted before use:

  • Platform Limitation: Supports Windows only. eyes_ocr depends on the built-in Windows OCR engine, so other platforms cannot provide the “read text” capability.
  • GUI Dependency: A browser GUI must be open. eyes_render normal rendering waits for approximately 15 seconds, the first Mermaid rendering takes approximately 45 seconds, and eyes_paste waits for pasting for approximately 60 seconds. If the GUI is not open, the tools will report errors.
  • Mermaid Loading: Mermaid diagrams rely on on-demand loading. On first rendering, the browser fetches the Mermaid library from a CDN and caches it after success. It can still be used without Mermaid installed locally, but the first use requires an internet connection.
  • OCR Results: Recognition results depend on Windows language packs. OCR performance for mixed Chinese-English text is average.
  • File Format: eyes_analyze supports only 8-bit RGB/RGBA non-interlaced PNG files.
  • Service Dependency: The attachments service is required.
  • Native Image Sending: The DSH input box natively supports pasting images, but sending images requires switching to a model that supports images (such as dsh-vision-router). dsh-omni-vision provides a local offline image reading path that is not subject to stream control limitations.

References

  • Plugin directory: https://www.skillhub.cn/plugins/Renji004/dsh-omni-vision
  • Source repository: https://github.com/Renji004/dsh-omni