Introduction

The core idea of DeepSeek Harness (DSH) is “everything is a plugin,” extending functionality through modular composition. However, many flagship models (such as deepseek-v4-pro) currently support text input only. Under the default configuration, if a user sends a message with a pasted image, the host rejects it directly.

The dsh-image-auto-describe plugin solves this problem by intercepting image messages and automatically converting them. It turns the image content into text and then passes that text to DeepSeek for processing, allowing a text-only model to “understand” images.

Plugin Overview

Plugin name: oldhan2423/dsh-image-auto-describe

Core value: It lets DeepSeek—an LLM that cannot natively “see” images—understand the images you send. When you paste an image into the conversation, the plugin first uses a vision model to read the image content into text, then passes that text to DeepSeek to continue answering—fully automatically, with no extra actions required from you.

License: MIT

Installation and Enablement

The plugin is installed through DSH’s plugin manager.

  1. Run the following command to add the plugin:
dsh plugin --profile web add github:oldHan2423/dsh-image-auto-describe
  1. Install the host patch (required).
    The official apiproxy currently exposes no public image admission interface, and the “reject image” logic is hardcoded. The plugin repository includes an idempotent patch to modify this logic. Run the following command to apply the patch (repeated execution is harmless):
node scripts/patch-seam.mjs

Configuration Requirements

Before using the plugin, the following conditions must be met:

  • DeepSeek Harness version: Requires 0.1.0-rc.6.
  • Vision model provider: A visual routing provider must be configured in the composed llm capability (for example, siliconflow or zhipu). By default, siliconflow is expected to provide Qwen/Qwen3-VL-32B-Instruct, and zhipu is expected to provide glm-4v-flash.

Configuration Items

In the Web plugin settings page, you can configure the following parameters:

  • candidates: A list of visual routing providers. They are attempted in order, and the first successful one takes effect. The type is { provider: string, model: string }[].
  • maxTokens: The token budget for each transcription. The default is 4096.

If the Harness version’s settings page does not yet support this, you can manually override the configuration in the profile’s cordis.patch.yml:

- id: image-auto-describe
  config:
    candidates:
      - provider: siliconflow
        model: Qwen/Qwen3-VL-32B-Instruct
      - provider: zhipu
        model: glm-4v-flash
    maxTokens: 4096

Usage

  1. Automatic recognition: Paste an image directly into the conversation. The plugin intercepts the message and calls a vision model to recognize it.
  2. Status display: During recognition, a status line saying “Recognizing image…” appears at the end of the conversation flow.
  3. Context injection: After recognition is complete, the text description is injected into DeepSeek’s context. DeepSeek only sees this text description and cannot access the image bytes.
  4. Original image retention: The original image is retained in the message (marked as presentationOnly). You can click to preview it, but it is not included in the model request.

Behavioral Details

  • Failure handling: If no routing provider is configured or all routes fail (for example, due to a missing API Key or insufficient quota), the message is rejected as-is. The plugin does not fabricate content.
  • Recognized content: The recognition result indicates which vision model produced it. Image details that the vision model does not transcribe cannot be answered by the chat model (even though the original image can still be previewed in the conversation).

Summary

This plugin intercepts image input through a patch layer, converts it into text, and then passes it to a text model for processing. It is suitable for most scenarios where screenshots, logs, or tables need to be reviewed. Users should note that DeepSeek always sees a text description rather than the image itself.

For more details, visit: GitHub repository