DeepSeek Harness (DSH) is a pluggable system. Pure text large language models (such as the DeepSeek V4 series) cannot directly process image input, requiring additional manual assistance for tasks involving charts, document screenshots, or interface screenshots. The dsh-image-describer plugin intercepts DSH’s message stream and transcribes image content into structured text, giving pure text models the ability to “look at images.”

Core Features

The plugin provides the following capabilities:

  • Dual-mode visual interaction: Supports direct pasting/drag-and-drop of images into the chat box (attachment mode) and proactive tool invocation in conversations (tool mode).
  • Automatic transcription and OCR: Automatically extracts text content (OCR) and layout structure from images, generating structured context.
  • Tool invocation support: The model can proactively trigger the describe_image tool during a conversation to perform targeted analysis on a local image path.
  • Context optimization: Converts image input into explicit system prompts, preventing pure text models from crashing when receiving binary data.

Installation

In the directory where the plugin is located, run the following command to add it to the Web Profile:

npx @deepseek-ai/dsh plugin --profile web add ./dsh-image-describer

After installation, start the DeepSeek Harness Web service:

dsh web --profile web

Usage

The plugin supports two complementary visual interaction scenarios.

Scenario 1: Pasting an Image in the Chat Box (Attachment Mode)

In the Web UI chat input box, directly use Ctrl + V to paste a screenshot, or drag an image file into the input box to send it.

  • The image is attached to the user message as an ImageBlock.
  • The plugin intercepts the request and invokes the configured multimodal vision model (default: MiniMax-M3) for OCR and content recognition.
  • The plugin replaces the image content with a system prompt containing the recognition results.
  • The pure text model answers directly based on the recognized text content, without performing additional file searches or tool calls.

Scenario 2: Mentioning a File Path in the Conversation (Tool Mode)

Mention a local image path in the conversation, and the model will proactively invoke the describe_image tool to analyze it.

  • User example: Analyze the microservice architecture in assets/architecture.png.
  • User example: What is the error code in the bottom-left corner of Screenshot.png?
  • After the model triggers the tool, the tool reads the local file and invokes the vision model for analysis, then returns the results to the model for the final answer.

Configuration

To use this plugin, complete the following two configuration steps.

1. Enable the Image Input Channel for Text Models

The DSH Web UI and API gateway check whether the model declares the image modality. Add the declaration for the corresponding text model in ~/.dsh/settings.yaml:

llm-pi-ai:
  providers:
    minimax-cn:
      apiKeyEnv: MINIMAX_CN_API_KEY
    opencode-go:
      apiKeyEnv: OPENCODE_GO_API_KEY
      modelOverrides:
        deepseek-v4-flash:
          input: [text, image]
        deepseek-v4-pro:
          input: [text, image]

2. Plugin Configuration

Configure the plugin parameters in the Profile’s cordis.patch.yml:

- insert:
    - id: image-describer-tool
      name: dsh-image-describer
      inject: [tools, llm, attachments]
      config:
        provider: minimax-cn       # 多模态模型提供者路由
        model: MiniMax-M3          # 视觉模型 ID
        maxTokens: 2048            # 单次分析最大输出 tokens
        timeoutMs: 60000           # 分析超时时间(毫秒)

Notes

  • Prerequisites: You must correctly configure the multimodal vision model (default: MiniMax-M3) and its API Key in settings.yaml; otherwise, the plugin will not work.
  • Model Declaration: You must declare input: [text, image] for the text model; otherwise, the image input feature cannot be enabled.
  • Performance Limits: The maximum output tokens for a single analysis are limited to 2048, and the timeout is set to 60000 milliseconds.