DeepSeek Harness (DSH) follows the “everything is a plugin” philosophy. When building agents, enabling the model to understand interface screenshots or image content is a common requirement. Directly integrating a vision model often requires additional configuration. The dsh-vision plugin registers the image_analyze tool and sends local images or URLs to an OpenAI-compatible vision model (defaulting to Google’s Gemini), enabling the main model, which originally supports only text input, to “see” images.
Plugin Overview¶
This is a model-inference plugin maintained by joyiok. It supports png/jpeg/gif/webp formats, and the default model is google/gemini-3.7-flash.
Core Features¶
- Provides the image_analyze tool: The agent can invoke this tool to analyze images.
- Multi-source input: Supports local paths, HTTP(S) URLs, and Data URLs.
- Secure handling: After an image is read, it is converted to Base64 in memory only and not written to disk.
- Automatic retries: When encountering a 429 or 5xx error, it automatically retries 2 times.
- Format support: Supports png, jpeg, gif, and webp formats.
Installation and Enablement¶
Install it using the official command. After installation, the dsh web process must be restarted.
dsh plugin --profile web add github:joyiok/dsh-vision
API Key Configuration¶
After installation, the API key must be configured. It can be set through environment variables, and the plugin resolves it via the dsh credentials service.
export OPENROUTER_API_KEY=sk-or-v1-xxxx
# 或
export VISION_API_KEY=sk-xxxx
Configuration Overrides¶
Configuration items can be overridden in the profile’s cordis.patch.yml. To disable the plugin, set disabled: true.
- id: vision
config:
apiKeyEnv: OPENROUTER_API_KEY
model: google/gemini-3.7-flash
baseUrl: https://openrouter.ai/api/v1/chat/completions
maxTokens: 2000
timeoutMs: 180000
maxImages: 4
maxImageBytes: 20971520
Typical Usage¶
The agent calls the image_analyze tool, passing the image path and the analysis question.
Agent:调用 image_analyze(images: ["/home/joy/Documents/memory-panel-preview.png"],
question: "检查这张界面截图的视觉设计和可用性问题")
Tool parameters:
* images: Required. An array, with a maximum of 4 images (local paths, URLs, or data URLs).
* question: Optional. The analysis question. By default, it performs detailed description and UI checks.
* model: Optional. Overrides the default vision model.
* maxTokens: Optional. Defaults to 2000.
Notes¶
- Runtime permissions: The plugin runs with the permissions of the current
dshprocess, so it is permitted to access local image paths allowed by the configuration file. - Request cancellation: Timeouts use
AbortSignal, and requests can be canceled by the upper layer. - Source code confirmation: Before use, please confirm the plugin source code and license (MIT).
With this plugin, agents can easily integrate vision capabilities for scenarios such as UI automated testing and interface analysis. The source code and documentation are available on GitHub.