DeepSeek Harness (DSH) follows the “everything is a plugin” philosophy. When building agents, enabling the model to understand interface screenshots or image content is a common requirement. Directly integrating a vision model often requires additional configuration. The dsh-vision plugin registers the image_analyze tool and sends local images or URLs to an OpenAI-compatible vision model (defaulting to Google’s Gemini), enabling the main model, which originally supports only text input, to “see” images.

Plugin Overview

This is a model-inference plugin maintained by joyiok. It supports png/jpeg/gif/webp formats, and the default model is google/gemini-3.7-flash.

Core Features

  • Provides the image_analyze tool: The agent can invoke this tool to analyze images.
  • Multi-source input: Supports local paths, HTTP(S) URLs, and Data URLs.
  • Secure handling: After an image is read, it is converted to Base64 in memory only and not written to disk.
  • Automatic retries: When encountering a 429 or 5xx error, it automatically retries 2 times.
  • Format support: Supports png, jpeg, gif, and webp formats.

Installation and Enablement

Install it using the official command. After installation, the dsh web process must be restarted.

dsh plugin --profile web add github:joyiok/dsh-vision

API Key Configuration

After installation, the API key must be configured. It can be set through environment variables, and the plugin resolves it via the dsh credentials service.

export OPENROUTER_API_KEY=sk-or-v1-xxxx
# 或
export VISION_API_KEY=sk-xxxx

Configuration Overrides

Configuration items can be overridden in the profile’s cordis.patch.yml. To disable the plugin, set disabled: true.

- id: vision
  config:
    apiKeyEnv: OPENROUTER_API_KEY
    model: google/gemini-3.7-flash
    baseUrl: https://openrouter.ai/api/v1/chat/completions
    maxTokens: 2000
    timeoutMs: 180000
    maxImages: 4
    maxImageBytes: 20971520

Typical Usage

The agent calls the image_analyze tool, passing the image path and the analysis question.

Agent:调用 image_analyze(images: ["/home/joy/Documents/memory-panel-preview.png"],
       question: "检查这张界面截图的视觉设计和可用性问题")

Tool parameters:
* images: Required. An array, with a maximum of 4 images (local paths, URLs, or data URLs).
* question: Optional. The analysis question. By default, it performs detailed description and UI checks.
* model: Optional. Overrides the default vision model.
* maxTokens: Optional. Defaults to 2000.

Notes

  • Runtime permissions: The plugin runs with the permissions of the current dsh process, so it is permitted to access local image paths allowed by the configuration file.
  • Request cancellation: Timeouts use AbortSignal, and requests can be canceled by the upper layer.
  • Source code confirmation: Before use, please confirm the plugin source code and license (MIT).

With this plugin, agents can easily integrate vision capabilities for scenarios such as UI automated testing and interface analysis. The source code and documentation are available on GitHub.