Introduction¶
The core design philosophy of DeepSeek Harness (DSH) is “everything is a plugin.” When developing agents, the main model often operates through a pure-text route (for example, DeepSeek V4 Flash routed via pi-ai), which means the model cannot directly handle image inputs. Existing community solutions usually require replacing the official adapter, an approach that is invasive and can easily break the stability of the plugin ecosystem. The dsh-image-vision plugin enables a pure-text main model to understand images by wrapping llm.resolveModelInfo and listening to the official agent/pre-step waterfall flow, without touching any host code or core files.
Core Features¶
- Generic image input: Supports pasting or dragging images into conversations. Even if the main model declares text-only support, images are sent to a vision model for description and ultimately returned to the main model as text.
- Feishu/Lark document support: Parses
<image token>elements in Feishu/Lark documents, automatically downloads and describes images, and injects the results into tool outputs. - Semantic understanding rather than OCR: Not only recognizes text, but also describes the semantics of scenes, people, tables, and charts.
- Local description tool: Provides the
describe_imagemodel tool, which can directly describe local image paths.
Installation and Activation¶
Install the plugin using the dsh plugin manager:
npx @deepseek-ai/dsh plugin --profile web add github:VeryInt/dsh-image-vision
After installation, restart DeepSeek Harness. The plugin will appear under Settings → Plugins.
Configuration¶
The plugin requires a vision model provider to be configured. Add a model configuration that supports image input in settings.yaml, and declare input: [text, image].
llm-pi-ai:
providers:
modelscope:
displayName: ModelScope
apiKeyEnv: MODELSCOPE_API_KEY
api: openai-completions
baseURL: https://api-inference.modelscope.cn/v1
models:
- id: Qwen/Qwen3-VL-8B-Instruct
name: Qwen3-VL-8B
input: [ text, image ]
Also, specify the vision model route in the plugin configuration (cordis.patch.yml). If both are left empty, the plugin automatically scans settings.yaml for the first model that declares input: [text, image].
Usage Examples¶
- Pasting images: Paste an image into the chat box; the model will output
[image content] .... - Feishu documents: After reading a document with a Feishu tool, the model can describe the images inside it.
- Local files: Call the
describe_imagetool and pass a local file path to obtain a description.
Notes¶
- Environment requirements: Requires DeepSeek Harness 0.1.0-rc.6 or later and Node.js >= 22.19.
- Model suffix: A vision model with the
VLsuffix (such asQwen/Qwen3-VL-8B) must be used. Pure-text models will not work even if image input is configured. - Ecosystem difference: Unlike community solutions that require replacing the official adapter, this plugin does not break the official plugin ecosystem and works with any pure-text route.
- Permission scope: The plugin runs with the permissions of the current DSH process. Before installation, review the source code and license.
Summary¶
dsh-image-vision is a lightweight DSH plugin. By wrapping existing llm service interfaces, it provides seamless image understanding capabilities for pure-text models. It does not modify host code or replace the official adapter, making it ideal for developers who need to add vision capabilities without disrupting their existing architecture.