Introduction¶
The plugin-based architecture of DeepSeek Harness (DSH) allows adding visual capabilities to text models. dsh-vision-tool is a plugin that provides vision routing for text-only models. It does this by automatically rewriting pasted images into content-addressed attachment references and invoking the Kimi vision model.
Core Features¶
The plugin works through two main components:
- vision-prompt: Intercepts
POST /api/session.prompt. When the active session model does not support image input, pasted images are persisted as content-addressed attachment references and rewritten in the text prompt as complete attachment reference JSON. Other requests (no image, or the model already supports images) are forwarded as-is, and the official/apitrust guardrails are reimplemented (DNS rebinding / cross-site defense). - vision-tool: Registers the global
analyze_imagetool. When the model invokes this tool, it passes an attachment reference (or a local file path); the tool routes the image to the Kimi vision model and returns a text description.
Installation and Activation¶
Use the following command to install the plugin (a SHA is required to pin the version):
dsh plugin --profile <name> add github:<you>/dsh-vision-tool#<sha>
After installing, restart the profile to make the configuration take effect:
dsh --profile <name>
Confirm that the plugin layer has been loaded (check whether the configuration output contains # == dsh-vision-tool):
dsh --profile <name> --dump-config
Typical Usage¶
- Paste an image: Paste an image into the chat box.
- Automatic rewriting:
vision-promptpersists the image to the Harness attachment store and rewrites the prompt to include the attachment reference. - Model invocation: The model invokes the registered
analyze_imagetool, passing an attachment reference or a local file path. - Configure the model: Switch the vision model by configuring
cordis.patch.yml(baseURL, model, apiKeyEnv).
Configuration and Dependencies¶
Required dependencies:
* dsh CLI
* KIMI_CODE_API_KEY (read from $DSH_HOME/.credentials.yaml or an environment variable)
* Dependency packages: @deepseek-ai/dsh-tools, @deepseek-ai/dsh-credentials, @deepseek-ai/dsh-host-apiproxy
Configuration example (in cordis.patch.yml):
- id: vision-tool
name: dsh-vision-tool
config:
baseURL: https://api.kimi.com/coding/v1
model: kimi-for-coding
apiKeyEnv: KIMI_CODE_API_KEY
maxImageBytes: 20971520
timeoutMs: 120000
Note:
kimi-for-codingonly acceptstemperature: 1(other values return HTTP 400). This plugin hard-codes that value by default and it cannot be configured.
Notes¶
- Security and storage: The request body limit is 160 MB. The plugin does not store prompt or image content; it only uses the Harness attachment store. Any failure falls back to passthrough.
- Logging: Operation logs are written to
$DSH_HOME/vision-trace.log. - Format support: png, jpg, jpeg, webp, and gif are supported.
Summary¶
dsh-vision-tool enables text models in DSH to invoke the Kimi vision model by rewriting prompts and registering a global tool. Before installing, review the source code and license (MIT).