DeepSeek Harness agents process text by default. When screenshot analysis or multi-image comparison is required, inserting images directly into the main session can quickly fill the context window, leading to degraded reasoning quality or wasted tokens. The dsh-vision-subagent plugin solves this by delegating image reading to an independent subagent. This subagent runs on a configured vision route (such as MiniMax or Kimi) and returns only the final text result, keeping image byte streams and intermediate context isolated from the main session.

Core Capabilities

This is a plugin that provides visual capabilities to DSH text agents. It allows agents to invoke a one-time subagent to analyze images through the vision_agent tool.

  • Context isolation: Image data streams and the intermediate reasoning of the vision model do not enter the main session window. Only the final analysis text is returned, preventing screenshots from polluting the main context.
  • Multi-round visual reasoning: The subagent can call read_image multiple times before answering to read files in the workspace, enabling more complex visual analysis.
  • Cost and routing separation: Vision calls are billed through the MiniMax/Kimi routing path, while the main model is only responsible for reasoning, avoiding expensive large-model token consumption on the same route.
  • Web paste support: When images are pasted or dragged into the web chat input, one visual analysis is automatically triggered, keeping the chat bubble interface clean.
  • Tool-based restoration: Provides the vision_image_fetch tool, which can restore analyzed images to the workspace with high fidelity for subsequent editing.

Installation and Configuration

The plugin is provided via an install package. Before installing, make sure the source path is ready. Run the following command to add the plugin to the current web configuration file:

dsh plugin --profile web add /path/to/dsh-vision-subagent

After installation, configure the vision route. Edit the ~/.dsh/profiles/web/cordis.patch.yml file and add the following configuration block:

- insert:
  - id: vision-subagent
    name: 'dsh-vision-subagent'
    config:
      provider: kimi-coding   # 或 minimax-cn / 自定义路由
      model: k3               # 或 MiniMax-M3 / MiniMax-VL-01

After configuration is complete, restart the dsh web service and start a new session to use it.

Usage Examples

In the web interface or through tool calls, the agent can directly ask to view a local image. For example:

“Look at ~/Desktop/error.png — what is the error?”

The agent will automatically invoke the subagent to analyze the image and return only a text description of the error in the main session.

Security and Limitations

Note the following when using the plugin:

  • Key management: The plugin does not manage keys directly. It obtains them through credential references from the configured vision route. The plugin itself does not access plaintext keys.
  • Image sources: The current version does not support remote image URLs (allowRemoteUrls: false). Images must be located within the session workspace (allowOutsideWorkspace: false). Local symbolic links are rejected.
  • Security limits: The subagent runtime is limited to a depth of 1 (maxDepth: 1). It is prohibited from delegating to other agents or modifying the filesystem, and is only allowed to read images.
  • Routing requirements: The provider and model fields must be configured manually; otherwise, the plugin remains dormant.

Summary

dsh-vision-subagent uses a subagent mechanism to provide lightweight visual analysis capabilities to text-based DSH agents, addressing context overflow and cost allocation issues. The project is maintained by ruby1304 and is released under the MIT license. The source code is available on GitHub.