Introduction

DeepSeek’s chat models currently do not support image input, and DeepSeek Harness (DSH) does not allow switching to a text-only model in sessions that contain images. The dsh-subagent-vision plugin solves this problem. It allows a text-only main agent to read images in a session without switching models.

Plugin Introduction

This is a DSH plugin maintained by niuniuaba. When the main agent needs vision capabilities, it delegates the task to a subagent, which is routed to a configured vision model. The main agent remains in a text-only state; only the subagent’s text result is merged back.

Core Features

  • subagent_vision delegation tool: a specific tool used to trigger the subagent.
  • Subagent read_image tool: the subagent uses this tool to read images.
  • /subagent-vision/paste route: a dedicated route for handling pasted image paths.
  • Native image ingestion in the Web GUI: pasting or dragging into the editor remains native.
  • Image-to-path conversion on send: when sending, images are converted to file paths so the main agent can pass them to the subagent.
  • Vision model selection: select the vision model via Settings > Vision Processing Model.

Installation and Enablement

Run the following command in the terminal to install the plugin:

dsh plugin --profile web add dsh-subagent-vision

After installation, restart dsh.

Typical Usage

  1. Paste an image into the editor, drag it into the editor, or provide a path/URL.
  2. Tell the agent the task (for example, “Read the pasted image…”).
  3. The plugin converts the image into a path and invokes the subagent’s read_image tool.
  4. The subagent reads the image and returns text.

Notes

  • Dependencies: A DSH installation is required, with its base package mounting subagent, tool-fs, attachments, and Web.
  • Vision model: At least one vision model must be configured in Settings > Models, and that model must declare image input. The default vision model is qwen/qwen3.8-max.
  • Image size: Image size is limited to 25 MB.
  • Subagent approval: Subagent approval is fixed to “Never”.
  • Logging: Image blocks are retained in the subagent’s logs.

Summary

This plugin adds vision capabilities to text models without requiring session switching. See the directory and the GitHub repository for more information.