Introduction

When developing agents in DeepSeek Harness (DSH), handling multimodal input usually requires manual checks. This plugin solves that problem: it detects images in user messages and automatically switches routing to a vision model, while keeping plain-text requests on the original model. This avoids the resource waste of calling a vision model during text-only turns, and prevents failures caused by calling a text-only model during image turns.

Plugin Introduction

@blue/dsh-image-model-router is a DSH plugin maintained by Blue-2571. Its core function is: when a user message contains images, route the agent request to a vision model (default: deepseek-v4-flash-vision-exp); when the message contains only text, continue using the original model (such as deepseek-v4-pro). The plugin uses a WeakMap internally for state marking, ensuring that the state is valid only within the current agent and step and does not leak across agents or turns.

Installation and Enabling

The plugin is disabled by default. After installation, it does not immediately change the existing routing logic.

1. Install via NPM:

dsh plugin --profile web add @blue/dsh-image-model-router

2. Install from a local Git checkout (development mode):

git clone https://github.com/Blue-2571/dsh-image-model-router.git
cd dsh-image-model-router
dsh plugin --profile web add .

Configuration Description

The plugin does not take effect automatically after installation. You must manually configure and enable it in the corresponding profile’s cordis.patch.yml file.

Configuration example:

- id: dsh-image-model-router
  disabled: false
  config:
    provider: deepseek-official
    model: deepseek-v4-flash-vision-exp

Configuration options:
* provider: Target provider ID for vision requests. Default: deepseek-official.
* model: Vision model ID. Default: deepseek-v4-flash-vision-exp. You can change this to any image-capable model supported by the provider.

How It Works

The plugin intercepts requests in the agent execution flow and detects images:
1. When a user message contains image blocks, the plugin marks the current agent during the agent/pre-step stage.
2. During the agent/request stage, the plugin checks the marker. If the marker exists, it swaps the provider and model for the vision configuration and clears the marker after processing completes.
3. If the message does not contain images, the plugin skips the marker and swap, and uses the original model configuration.

Like DSH’s own image policy, the plugin traverses nested tool-result content, so images returned by tools are also treated as input images.

Notes

  • Disabled by default: After installation, you must explicitly set disabled to false in cordis.patch.yml to enable the plugin.
  • Apply changes with a refresh: After modifying the configuration, you must restart the profile (dsh web) and perform a hard refresh in the browser for the changes to take effect.
  • Environment requirements: Requires the Node.js environment ^22.19.0 || >=24.0.0.

Summary

This plugin reduces the complexity of handling multimodal input in DSH by automating the model routing logic. For more details and source code, refer to GitHub or the DSH plugin catalog.