In the DeepSeek Harness (DSH) ecosystem, model inference is often text-centric. To introduce visual capabilities into an Agent preset, you usually need to reconcile image upload with security policies. The dsh-tool-vision plugin solves this by providing the model-side tool image_describe, which calls DashScope’s OpenAI-compatible API so that the model can directly describe local images. It also provides a paste-bridge feature that converts pasted images into path text, avoiding triggering the host’s image admission restrictions.

1. What is this

This is a DSH plugin maintained by wanshichenguang and released under the MIT license. It is a model-oriented image description tool that uses the DashScope OpenAI-compatible API (default qwen3.7-flash model) to read local images and generate descriptions. It also provides a paste-bridge feature that converts pasted images into path text, avoiding triggering the host’s image admission restrictions.

2. Installation and Enablement

To install this plugin, use DSH’s plugin command. Run the following command under the Web preset to complete installation:

dsh plugin --profile web add dsh-tool-vision

After installation, restart the dsh web process. The tool is registered as image_describe and appears in the model’s tool directory.

3. Core features

  • image_describe tool: This is the core feature. The invocation is image_describe(path, prompt). path is the image path (supports absolute or workspace-relative paths), and prompt is an optional description instruction. If prompt is not provided, a default instruction is used to describe the image in detail.
  • Pasted image bridge: In the Web UI, if the session’s model is confirmed to be a text-only model, pasting or dragging an image uploads the image to the plugin’s host route, releases the draft, and appends [image: <path>] text to the prompt. This both bypasses the host’s image admission check and allows the model to read the file through the tool. For models with visual capabilities or unconfirmed models, images remain natively inline.
  • Configurable API Key: The plugin does not include any credentials; developers must configure them themselves. It supports specifying via environment variables or a configuration file.

4. Configuration and usage

The plugin depends on the DashScope API and requires a configured Key.

Set an environment variable:

DASHSCOPE_API_KEY=sk-xxxx

Specify via a configuration file:
Add the configuration to cordis.patch.yml:

- id: tool-vision
  name: 'dsh-tool-vision'
  config:
    apiKey: sk-xxxx

Tool call example:
When generating a response, the model calls this tool if it needs to analyze an image. For example, the model may issue the following call request:

{
  "tool_calls": [{
    "function": {
      "name": "image_describe",
      "arguments": "{\"path\": \"/workspace/image.png\", \"prompt\": \"这是什么?\"}"
    }
  }]
}

Key configuration options:
* baseUrl: DashScope OpenAI-compatible endpoint; default is https://dashscope.aliyuncs.com/compatible-mode/v1.
* model: Visual model used; default is qwen3.7-flash.
* defaultPrompt: Default instruction when prompt is not provided.
* maxImageBytes: Maximum size for a single image; default is 20MB.

5. Notes

  • Credential security: The plugin package does not contain any credentials; the Key only exists in your environment variables or configuration file.
  • HTTP behavior: The HTTP client sets redirect: 'error'; the configured endpoint must directly accept credentials to prevent Key leakage caused by redirects.
  • API limitation: Currently, only DashScope OpenAI-compatible endpoints are supported.
  • Permission policy: The tool does not request ctx.approval when executed. If the deployment environment requires a confirmation mechanism, you must additionally configure a tools/pre-execute policy at deployment time.

dsh-tool-vision provides DSH users with a low-intrusion way to add visual capabilities to text models. By using a model-side tool, it decouples image analysis from prompt generation, making it suitable for scenarios that require handling local visual context within the conversation flow.

Plugin catalog: catalog_url
Source code: github_url