Introduction

The plugin-based architecture of DeepSeek Harness (DSH) allows adding visual capabilities to text models. dsh-vision-tool is a plugin that provides vision routing for text-only models. It does this by automatically rewriting pasted images into content-addressed attachment references and invoking the Kimi vision model.

Core Features

The plugin works through two main components:

  1. vision-prompt: Intercepts POST /api/session.prompt. When the active session model does not support image input, pasted images are persisted as content-addressed attachment references and rewritten in the text prompt as complete attachment reference JSON. Other requests (no image, or the model already supports images) are forwarded as-is, and the official /api trust guardrails are reimplemented (DNS rebinding / cross-site defense).
  2. vision-tool: Registers the global analyze_image tool. When the model invokes this tool, it passes an attachment reference (or a local file path); the tool routes the image to the Kimi vision model and returns a text description.

Installation and Activation

Use the following command to install the plugin (a SHA is required to pin the version):

dsh plugin --profile <name> add github:<you>/dsh-vision-tool#<sha>

After installing, restart the profile to make the configuration take effect:

dsh --profile <name>

Confirm that the plugin layer has been loaded (check whether the configuration output contains # == dsh-vision-tool):

dsh --profile <name> --dump-config

Typical Usage

  1. Paste an image: Paste an image into the chat box.
  2. Automatic rewriting: vision-prompt persists the image to the Harness attachment store and rewrites the prompt to include the attachment reference.
  3. Model invocation: The model invokes the registered analyze_image tool, passing an attachment reference or a local file path.
  4. Configure the model: Switch the vision model by configuring cordis.patch.yml (baseURL, model, apiKeyEnv).

Configuration and Dependencies

Required dependencies:
* dsh CLI
* KIMI_CODE_API_KEY (read from $DSH_HOME/.credentials.yaml or an environment variable)
* Dependency packages: @deepseek-ai/dsh-tools, @deepseek-ai/dsh-credentials, @deepseek-ai/dsh-host-apiproxy

Configuration example (in cordis.patch.yml):

- id: vision-tool
  name: dsh-vision-tool
  config:
    baseURL: https://api.kimi.com/coding/v1
    model: kimi-for-coding
    apiKeyEnv: KIMI_CODE_API_KEY
    maxImageBytes: 20971520
    timeoutMs: 120000

Note: kimi-for-coding only accepts temperature: 1 (other values return HTTP 400). This plugin hard-codes that value by default and it cannot be configured.

Notes

  • Security and storage: The request body limit is 160 MB. The plugin does not store prompt or image content; it only uses the Harness attachment store. Any failure falls back to passthrough.
  • Logging: Operation logs are written to $DSH_HOME/vision-trace.log.
  • Format support: png, jpg, jpeg, webp, and gif are supported.

Summary

dsh-vision-tool enables text models in DSH to invoke the Kimi vision model by rewriting prompts and registering a global tool. Before installing, review the source code and license (MIT).

Plugin Directory
GitHub Repository