Introduction

When using a text-only model such as DeepSeek as the primary model in DeepSeek Harness (DSH), there is one unavoidable limitation: by default, Harness does not allow sending messages with images to sessions where the current model does not support images. If you want the model to see an error screenshot, a UI layout, or a photo of a table, you either have to convert the image content to text first and send that, or switch to a multimodal primary model.

dsh-eyes offers an alternative: images are pasted and sent as usual and persisted in the background; the primary model invokes the view_image tool whenever it needs to look at an image, and the recognition work is performed by any OpenAI-compatible vision endpoint (default: Alibaba Cloud Bailian Qwen). From a usage perspective, this is close to the model natively having multimodality. The following sections introduce its features, installation, and usage.

What is this

dsh-eyes is a plugin for DeepSeek Harness, authored by Leeminjing, released under the MIT license, and classified under “Model Reasoning” in the community catalog. It solves a very concrete problem: allowing messages with images to pass send admission in sessions with a text-only primary model, being persisted, and being viewable on demand by the model.

DSH’s philosophy is “everything is a plugin,” with capability gaps filled by plugins. dsh-eyes belongs to this category: it does not modify the primary model or add a switch, but turns “image viewing” into a tool that the model can call at any time.

Core features

  • Send admission passthrough: allows messages with images to pass the send admission check and be persisted;
  • Strip images before dispatch: replaces image blocks in the request with indexed reference descriptions (【图片N attachment_id=…】), so the primary model receives text-only content;
  • view_image tool: supports viewing a single image by attachment_id, viewing multiple images at once with an attachment_ids array, or reading a local image with image_path;
  • OpenAI-compatible vision endpoint: supports both Chat Completions and Responses API, controlled by VISION_API_STYLE (auto / chat / responses);
  • Session isolation and persistence: image reference indices are sharded by sessionId and persisted to .dsh/attachments/v1/dsh-eyes-index.json, so historical images can still be viewed via attachment_id after a restart or context compression;
  • Automatic primary-model adaptation: any text-only primary model (regardless of provider) is automatically protected; natively multimodal primary models that already support images are not affected.

Installation and configuration

First install the plugin, then configure the environment variables for the vision endpoint, and finally restart dsh. Using Windows as an example, the full steps are as follows:

# 1) 安装插件(github 方式,也可以换成 npm 包名)
dsh plugin --profile web add github:Leeminjing/dsh-eyes

# 2) 配置 API Key(换成你所用视觉提供商的 key)
setx VISION_API_KEY "sk-你的key"

# 3) 配置视觉模型(必填,无默认值)
setx VISION_MODEL "qwen-vl-plus"

# 4) 配置接口端点(默认百炼;换其他提供商时必改)
setx VISION_ENDPOINT "https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions"

# 5) 重启 dsh,让环境变量生效

The four environment variables mean the following:

Environment variable Description Default
VISION_API_KEY API key for the vision endpoint Required
VISION_MODEL Vision model name, e.g., qwen-vl-plus Required; no default
VISION_ENDPOINT Vision endpoint URL https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions
VISION_API_STYLE API style auto

Note that setx only takes effect for newly started processes, so dsh must be restarted after configuration. When VISION_API_STYLE is auto, the style is inferred automatically from the endpoint path: if the path contains chat/completions, Chat Completions is used; if it contains responses, the Responses API is used; a bare base URL defaults to Chat Completions and the path is completed automatically. You can also force chat or responses; the plugin normalizes the endpoint path to the corresponding protocol.

Common vision providers

Provider Endpoint Model examples
Alibaba Cloud Bailian https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions qwen-vl-plus / qwen-vl-max
OpenAI https://api.openai.com/v1/chat/completions gpt-4o / gpt-4o-mini
Moonshot https://api.moonshot.cn/v1/chat/completions moonshot-v1-8k-vision-preview
OpenRouter https://openrouter.ai/api/v1/chat/completions qwen/qwen2.5-vl-72b-instruct

The native APIs of Anthropic and Gemini are not directly supported and require their respective OpenAI-compatible gateways.

Override configuration with cordis.patch.yml

In addition to environment variables, you can also pass a config value for that entry in cordis.patch.yml, which overrides the defaults and environment variables:

- insert:
    - id: dsh-eyes
      name: dsh-eyes
      config:
        apiKey: sk-xxx          # 同 VISION_API_KEY
        model: qwen-vl-plus     # 同 VISION_MODEL
        # endpoint, targetProvider, maxImageBytes 同理

The target primary model is determined by targetProvider in the code, defaulting to deepseek-official; maxImageBytes controls the maximum local image size, defaulting to 15 MB.

Typical usage

No switch is required, and no manual tool invocation is needed:

  1. In a session, paste a screenshot directly with Ctrl+V, or drag and drop / use the attachment option to add it;
  2. Ask as if sending a normal message, for example, “What does this image say?”;
  3. The primary model receives the image reference description and decides on its own during reasoning to call view_image(attachment_id=…);
  4. The vision model extracts a description or OCR text, and the primary model continues answering based on that text.

The image remains in the background. In any subsequent turn, you can continue asking about the same image, such as “Look at the number on the second line again.” Because the reference index is persisted, even if the context is compressed or the dsh process is restarted, the model can still view historical images through attachment_id.

Working principle

The plugin intervenes at three points:

  1. Admission passthrough: It wraps llm.resolveModelInfo, causing the target primary model to declare inputModalities: ['text','image']. As a result, messages with images pass the Host’s send admission check and are persisted, generating an attachment_id.
  2. Image stripping: It wraps llm.streamWithRegistration. Before the request is dispatched to the model, it replaces each image block (including those nested inside tool-result) with a 【图片N attachment_id=…】 reference description and registers a sessionId-sharded attachment_id → ImageAttachmentRef mapping.
  3. view_image invocation: When the primary model needs to view an image, the plugin reads the image bytes by attachment_id (or attachment_ids, image_path), converts them into a data: URL, POSTs them to the configured vision endpoint, and passes the returned text to the primary model for answering. When viewing multiple images at once, the returned text is segmented by 【图片N】.

These two methods are wrapped directly because Harness currently has no public extension points for “image capability judgment in send admission” and “stripping images before dispatch.”

Applicable scenarios and notes

Suitable scenario: DSH users whose primary model must be DeepSeek or another text-only model, but who occasionally need image viewing (screenshot OCR, UI inspection, table recognition, etc.). If your primary model is already multimodal, this plugin will not intervene and there is no need to install it.

Before installing and using it, note the following:

  • The plugin runs with the permissions of the current dsh process. Before installing, it is recommended to read the source code and the MIT license first, and proceed only after confirming it is acceptable;
  • Manage the API key with environment variables or a credential service; do not hard-code it into a repository;
  • There are two independent and unrelated size limits: the local image_path image-reading limit is 15 MB (maxImageBytes); pasted or attachment images are subject to the Harness attachment-storage limit (default 5 MB);
  • A text-only primary model will be displayed as “supports images” in the model selector. This is intentional, to allow messages with images to pass, and does not mean the model itself has become multimodal;
  • The vision endpoint must be OpenAI-compatible (Chat Completions or Responses API).

Summary

The dsh-eyes approach separates “image viewing” from the primary model’s capability requirements: images remain in the background, the model calls view_image on demand, and recognition is delegated to an external OpenAI-compatible vision endpoint. For DSH users blocked by “a text-only model cannot send images,” this is a minimal-change, immediately usable solution.

  • GitHub: https://github.com/Leeminjing/dsh-eyes
  • Community catalog: https://www.skillhub.cn/plugins/Leeminjing/dsh-eyes

It should be noted that the community catalog is an independent site and has no official affiliation with DeepSeek or High-Flyer; the GitHub repository is authoritative for plugin information.