Introduction¶
DeepSeek’s chat models currently do not support image input, and DeepSeek Harness (DSH) does not allow switching to a text-only model in sessions that contain images. The dsh-subagent-vision plugin solves this problem. It allows a text-only main agent to read images in a session without switching models.
Plugin Introduction¶
This is a DSH plugin maintained by niuniuaba. When the main agent needs vision capabilities, it delegates the task to a subagent, which is routed to a configured vision model. The main agent remains in a text-only state; only the subagent’s text result is merged back.
Core Features¶
subagent_visiondelegation tool: a specific tool used to trigger the subagent.- Subagent
read_imagetool: the subagent uses this tool to read images. /subagent-vision/pasteroute: a dedicated route for handling pasted image paths.- Native image ingestion in the Web GUI: pasting or dragging into the editor remains native.
- Image-to-path conversion on send: when sending, images are converted to file paths so the main agent can pass them to the subagent.
- Vision model selection: select the vision model via
Settings > Vision Processing Model.
Installation and Enablement¶
Run the following command in the terminal to install the plugin:
dsh plugin --profile web add dsh-subagent-vision
After installation, restart dsh.
Typical Usage¶
- Paste an image into the editor, drag it into the editor, or provide a path/URL.
- Tell the agent the task (for example, “Read the pasted image…”).
- The plugin converts the image into a path and invokes the subagent’s
read_imagetool. - The subagent reads the image and returns text.
Notes¶
- Dependencies: A DSH installation is required, with its base package mounting subagent, tool-fs, attachments, and Web.
- Vision model: At least one vision model must be configured in
Settings > Models, and that model must declare image input. The default vision model isqwen/qwen3.8-max. - Image size: Image size is limited to 25 MB.
- Subagent approval: Subagent approval is fixed to “Never”.
- Logging: Image blocks are retained in the subagent’s logs.
Summary¶
This plugin adds vision capabilities to text models without requiring session switching. See the directory and the GitHub repository for more information.