Introduction¶
DeepSeek Harness (DSH) agents perform well when handling text, but they are often helpless when it comes to images. When users upload images or reference screenshots, a text-only agent cannot directly perceive visual information. This plugin works by registering a system skill, enabling the agent to call vision CLI tools when it encounters an image, converting image content into text descriptions or coordinate information, and thereby addressing the lack of visual perception.
Plugin Overview¶
- Name: dsh-plugin-vision-toolkit
- Author: YYTbit
- Category: Model Inference
- License: MIT
Core Features¶
This plugin provides four basic tools designed to help agents understand image content:
* glance: describe an image, ask questions about it, or perform OCR.
* ground: locate a specific element in an image and return its bounding box coordinates.
* detect: find all instances of a specific type of element in an image.
* crop: crop a specific region from an image based on coordinates.
Installation and Configuration¶
Installing the plugin requires specifying a profile configuration:
dsh plugin --profile your-profile add dsh-plugin-vision-toolkit
After installation, configure environment variables to connect to the vision API:
export VISION_API_KEY=sk-xxx # Vision API key (falls back to DEEPSEEK_API_KEY if not set)
export VISION_BASE_URL=https://... # API endpoint (falls back to DEEPSEEK_BASE_URL if not set)
export VISION_MODEL=deepseek-vl2 # Vision model name
Usage Examples¶
The following are verified use cases:
- Describe an image:
glance screenshot.png
- Ask a question about an image:
glance screenshot.png -q "What error is shown?"
- OCR recognition:
glance screenshot.png --ocr
- Locate an element:
ground screenshot.png "the login button"
# Output example: 450,820,620,870
- Detect all instances:
detect screenshot.png "buttons"
- Crop a region:
crop screenshot.png 450,820,620,870 button.png
Dependencies and Notes¶
- Dependencies: The plugin depends on
@deepseek-ai/cordisand the image processing librarysharp. - Runtime permissions: The plugin runs with the permissions of the current DSH process. It is recommended to review the source code and license before installation.
- Data security: The agent does not directly access raw pixels; instead, it obtains text descriptions or coordinate data through tools.
Summary¶
By registering skills, this tool adds visual capabilities to DSH agents. Developers only need to configure the environment variables to enable agents to process screenshots, perform OCR, or locate UI elements, thereby expanding the application boundaries of agents.