Introduction¶
When using a text-only model such as DeepSeek as the primary model in DeepSeek Harness (DSH), there is one unavoidable limitation: by default, Harness does not allow sending messages with images to sessions where the current model does not support images. If you want the model to see an error screenshot, a UI layout, or a photo of a table, you either have to convert the image content to text first and send that, or switch to a multimodal primary model.
dsh-eyes offers an alternative: images are pasted and sent as usual and persisted in the background; the primary model invokes the view_image tool whenever it needs to look at an image, and the recognition work is performed by any OpenAI-compatible vision endpoint (default: Alibaba Cloud Bailian Qwen). From a usage perspective, this is close to the model natively having multimodality. The following sections introduce its features, installation, and usage.
What is this¶
dsh-eyes is a plugin for DeepSeek Harness, authored by Leeminjing, released under the MIT license, and classified under “Model Reasoning” in the community catalog. It solves a very concrete problem: allowing messages with images to pass send admission in sessions with a text-only primary model, being persisted, and being viewable on demand by the model.
DSH’s philosophy is “everything is a plugin,” with capability gaps filled by plugins. dsh-eyes belongs to this category: it does not modify the primary model or add a switch, but turns “image viewing” into a tool that the model can call at any time.
Core features¶
- Send admission passthrough: allows messages with images to pass the send admission check and be persisted;
- Strip images before dispatch: replaces image blocks in the request with indexed reference descriptions (
【图片N attachment_id=…】), so the primary model receives text-only content; view_imagetool: supports viewing a single image byattachment_id, viewing multiple images at once with anattachment_idsarray, or reading a local image withimage_path;- OpenAI-compatible vision endpoint: supports both Chat Completions and Responses API, controlled by
VISION_API_STYLE(auto/chat/responses); - Session isolation and persistence: image reference indices are sharded by
sessionIdand persisted to.dsh/attachments/v1/dsh-eyes-index.json, so historical images can still be viewed viaattachment_idafter a restart or context compression; - Automatic primary-model adaptation: any text-only primary model (regardless of provider) is automatically protected; natively multimodal primary models that already support images are not affected.
Installation and configuration¶
First install the plugin, then configure the environment variables for the vision endpoint, and finally restart dsh. Using Windows as an example, the full steps are as follows:
# 1) 安装插件(github 方式,也可以换成 npm 包名)
dsh plugin --profile web add github:Leeminjing/dsh-eyes
# 2) 配置 API Key(换成你所用视觉提供商的 key)
setx VISION_API_KEY "sk-你的key"
# 3) 配置视觉模型(必填,无默认值)
setx VISION_MODEL "qwen-vl-plus"
# 4) 配置接口端点(默认百炼;换其他提供商时必改)
setx VISION_ENDPOINT "https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions"
# 5) 重启 dsh,让环境变量生效
The four environment variables mean the following:
| Environment variable | Description | Default |
|---|---|---|
VISION_API_KEY |
API key for the vision endpoint | Required |
VISION_MODEL |
Vision model name, e.g., qwen-vl-plus |
Required; no default |
VISION_ENDPOINT |
Vision endpoint URL | https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions |
VISION_API_STYLE |
API style | auto |
Note that setx only takes effect for newly started processes, so dsh must be restarted after configuration. When VISION_API_STYLE is auto, the style is inferred automatically from the endpoint path: if the path contains chat/completions, Chat Completions is used; if it contains responses, the Responses API is used; a bare base URL defaults to Chat Completions and the path is completed automatically. You can also force chat or responses; the plugin normalizes the endpoint path to the corresponding protocol.
Common vision providers¶
| Provider | Endpoint | Model examples |
|---|---|---|
| Alibaba Cloud Bailian | https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions |
qwen-vl-plus / qwen-vl-max |
| OpenAI | https://api.openai.com/v1/chat/completions |
gpt-4o / gpt-4o-mini |
| Moonshot | https://api.moonshot.cn/v1/chat/completions |
moonshot-v1-8k-vision-preview |
| OpenRouter | https://openrouter.ai/api/v1/chat/completions |
qwen/qwen2.5-vl-72b-instruct |
The native APIs of Anthropic and Gemini are not directly supported and require their respective OpenAI-compatible gateways.
Override configuration with cordis.patch.yml¶
In addition to environment variables, you can also pass a config value for that entry in cordis.patch.yml, which overrides the defaults and environment variables:
- insert:
- id: dsh-eyes
name: dsh-eyes
config:
apiKey: sk-xxx # 同 VISION_API_KEY
model: qwen-vl-plus # 同 VISION_MODEL
# endpoint, targetProvider, maxImageBytes 同理
The target primary model is determined by targetProvider in the code, defaulting to deepseek-official; maxImageBytes controls the maximum local image size, defaulting to 15 MB.
Typical usage¶
No switch is required, and no manual tool invocation is needed:
- In a session, paste a screenshot directly with
Ctrl+V, or drag and drop / use the attachment option to add it; - Ask as if sending a normal message, for example, “What does this image say?”;
- The primary model receives the image reference description and decides on its own during reasoning to call
view_image(attachment_id=…); - The vision model extracts a description or OCR text, and the primary model continues answering based on that text.
The image remains in the background. In any subsequent turn, you can continue asking about the same image, such as “Look at the number on the second line again.” Because the reference index is persisted, even if the context is compressed or the dsh process is restarted, the model can still view historical images through attachment_id.
Working principle¶
The plugin intervenes at three points:
- Admission passthrough: It wraps
llm.resolveModelInfo, causing the target primary model to declareinputModalities: ['text','image']. As a result, messages with images pass the Host’s send admission check and are persisted, generating anattachment_id. - Image stripping: It wraps
llm.streamWithRegistration. Before the request is dispatched to the model, it replaces eachimageblock (including those nested insidetool-result) with a【图片N attachment_id=…】reference description and registers asessionId-shardedattachment_id→ImageAttachmentRefmapping. view_imageinvocation: When the primary model needs to view an image, the plugin reads the image bytes byattachment_id(orattachment_ids,image_path), converts them into adata:URL, POSTs them to the configured vision endpoint, and passes the returned text to the primary model for answering. When viewing multiple images at once, the returned text is segmented by【图片N】.
These two methods are wrapped directly because Harness currently has no public extension points for “image capability judgment in send admission” and “stripping images before dispatch.”
Applicable scenarios and notes¶
Suitable scenario: DSH users whose primary model must be DeepSeek or another text-only model, but who occasionally need image viewing (screenshot OCR, UI inspection, table recognition, etc.). If your primary model is already multimodal, this plugin will not intervene and there is no need to install it.
Before installing and using it, note the following:
- The plugin runs with the permissions of the current dsh process. Before installing, it is recommended to read the source code and the MIT license first, and proceed only after confirming it is acceptable;
- Manage the API key with environment variables or a credential service; do not hard-code it into a repository;
- There are two independent and unrelated size limits: the local
image_pathimage-reading limit is 15 MB (maxImageBytes); pasted or attachment images are subject to the Harness attachment-storage limit (default 5 MB); - A text-only primary model will be displayed as “supports images” in the model selector. This is intentional, to allow messages with images to pass, and does not mean the model itself has become multimodal;
- The vision endpoint must be OpenAI-compatible (Chat Completions or Responses API).
Summary¶
The dsh-eyes approach separates “image viewing” from the primary model’s capability requirements: images remain in the background, the model calls view_image on demand, and recognition is delegated to an external OpenAI-compatible vision endpoint. For DSH users blocked by “a text-only model cannot send images,” this is a minimal-change, immediately usable solution.
- GitHub: https://github.com/Leeminjing/dsh-eyes
- Community catalog: https://www.skillhub.cn/plugins/Leeminjing/dsh-eyes
It should be noted that the community catalog is an independent site and has no official affiliation with DeepSeek or High-Flyer; the GitHub repository is authoritative for plugin information.