Introduction

The text models in DeepSeek Harness (DSH), such as deepseek-v4-flash, cannot process images directly. To enable visual capabilities, you typically need to integrate a multimodal model or apply patches to the host. The analyze-image-tool plugin registers an analyze_image tool and provides a “visual bridging” layer for text models, allowing them to answer questions about images through any OpenAI-compatible vision/multimodal endpoint.

Plugin Overview

This is a general-purpose visual interface plugin maintained by CaseyTso. It is not bound to a specific provider and allows users to configure baseURL, apiKey, and model to connect to any endpoint that supports the OpenAI format, including SiliconFlow, DashScope, Ollama, and others.

Key Features

  1. Generic endpoint support: Supports any OpenAI-compatible endpoint, allowing users to switch model providers without modifying plugin code.
  2. Sandbox-safe reading: Local images are preferentially read through the host sandbox’s fs channel, with support for HTTP/HTTPS and Data URLs.
  3. Schema safety: Uses defineTool to compile parameters into standard JSON Schema, avoiding session 400 errors caused by a non-object root in the schema.
  4. Robustness design: Automatically cleans up API keys in error messages and automatically strips reasoning blocks from the model’s reasoning process.
  5. Structured response: Returns a complete structure containing text, model name, and token usage.
  6. Coexistence: The tool name analyze_image does not conflict with community tools (such as view_image).
  7. WebUI settings panel: Provides a visual configuration interface, with support for saving multiple configuration presets and connection testing.

Installation and Activation

Before installing, ensure the environment meets the requirements: DeepSeek Harness (dsh) version 0.1.x and Node.js version ≥ 18.17.

Use the following commands to install the plugin (installed from the main branch, including runtime pasted-image bridging):

dsh plugin --profile web add github:CaseyTso/dsh-analyze-image-tool#main
dsh --profile web

The plugin is automatically installed as a profile layer and does not require manually editing cordis.patch.yml.

Configuration Guide

It is recommended to use the WebUI panel for configuration. Click the eye icon in the upper-right corner of the Web conversation header, or go to the plugin configuration card in System Settings to edit it.

The main configuration items include:
* baseURL: The target endpoint address (such as SiliconFlow or a local Ollama instance).
* apiKey: The API key. If left empty, it is read from the chain.
* model: The ID of the multimodal model to use.
* maxImageBytes: The maximum size in bytes allowed for image uploads, default 10MB.
* defaultQuestion: The default description request when prompt is not specified.

Configuration example (cordis.patch.yml):

- id: analyze-image-tool
  config:
    baseURL: https://api.siliconflow.cn/v1
    apiKey: ***
    model: Qwen/Qwen3-VL-32B-Instruct
    maxTokens: 2048
    timeoutMs: 60000
    maxImageBytes: 10485760
    defaultQuestion: 描述这张图片的细节,包括可见文字和布局。

Usage Example

When the model calls the tool, pass an image path or attachment ID:

模型: 看下 ~/Desktop/error.png 是什么报错
模型 → analyze_image(path="~/Desktop/error.png", prompt="这个报错的完整文本是什么?")
     ← "TypeError: Cannot read properties of undefined (reading 'map') at …"
模型: 这是一个 … 建议 …

For text models, the plugin automatically handles the logic for pasted images and converts it into an analyze_image(attachment_id="...") call.

Notes

  1. Version differences: The v0.1.0 tag is limited to local paths and URLs and does not include the pasted-image bridging feature. Installing the #main branch includes that feature.
  2. Pasted image limitations: The attachment reference index for pasted images is only valid within the current process. After a server restart, old attachment IDs cannot be retrieved from historical sessions, and the image must be pasted again.
  3. Permissions and security: The plugin runs with the current DSH process permissions. Calling the endpoint generates API requests and token consumption, so configure keys and timeouts according to your actual needs.

Conclusion

analyze-image-tool solves the common problem of adding visual capabilities to DSH text models through a standardized tool interface and flexible endpoint configuration. It is suitable for scenarios where you need to extend multimodal capabilities without modifying host code. For more details, refer to the plugin directory and the GitHub repository.