Introduction

When developing agents based on DeepSeek Harness (DSH), we often encounter a specific problem: the project is configured with text-only models such as DeepSeek or GLM, but when handling requests that include images (such as screenshot analysis or chart reading), image rejection or errors occur. Manually switching routes can solve the problem, but it disrupts the agent’s continuity and established behavior.

As a DSH plugin, the core value of dsh-mmroute lies in solving this problem “transparently.” It does not require changes to existing model invocation habits. Instead, it intervenes in the llm/stream flow as an intermediate layer, intercepting, transcribing, and dispatching images in requests sent by text-only models, ensuring that the entire agent flow can handle visual tasks smoothly.

Plugin Positioning

  • Name: dsh-mmroute
  • Maintainer: jmxsxwyzjdwl
  • Category: Model inference
  • License: MIT
  • Core capability: Provides image modality scheduling for every model route in DSH, across the entire agent flow.

Core Capabilities

The plugin mainly handles the following scenarios and functions:

  1. End-to-end image modality scheduling
    From user-uploaded images and read_image results to rendered images produced by MCP tools (such as Figma screenshots), all images are handled uniformly.

  2. Automatic transcription for text-only models
    For text-only models (or models not marked as multimodal), the plugin transcribes images in their requests into detailed textual descriptions (including verbatim transcription, chart data, layout, and color information).

  3. Automatic recovery from image-related errors
    If a text-only model (or gateway) rejects images and causes the request to fail, the plugin automatically reroutes to an understanding model and retries until success or the retry limit is reached.

  4. Historical image replay with summaries
    When the same image appears again in the conversation history, it is replayed as a summary (≤1200 characters), preventing long sessions from bloating due to full resending of old images.

  5. Structured transcription and vision_relook
    Transcriptions include structured information such as image type classification, full text content, and chart data. When responding, text-only models can invoke the vision_relook tool to re-examine specific image details, forming a “command-and-execution” collaboration loop.

  6. Multi-format support
    Supports common image formats such as PNG, JPEG, WebP, GIF.

Installation and Enablement

The installation process is straightforward and can be completed using the official command:

dsh plugin --profile web add dsh-mmroute

After installation, restart the DSH Web process for the plugin to take effect. The plugin automatically writes dependencies into the profile’s package.json.

Usage

After enabling the plugin, the workflow is as follows:

  1. Open Settings → Multimodal Routing.
  2. In the “Multimodal Understanding Model” dropdown, select:
    • Automatic: Use the first discovered native multimodal model.
    • Specify: Manually select a configured multimodal model (it can be a free cloud model or a local model).
  3. (Optional) In “Model Modality Marking”, explicitly select Default / Multimodal / Text-only for individual models. This overrides the automatic detection result.
  4. In the conversation, directly paste or upload an image, or let the agent invoke MCP tools to generate images, and you will see the automatic transcription effect.

Note: Starting with v0.6.0, the settings page has been simplified and only retains the understanding model and modality marking. Features such as automatic routing, error recovery, and historical image summaries are enabled by default and run continuously in the background.

Configuration and Mechanism Details

  • Configuration persistence: The plugin’s configuration is stored in the $DSH_HOME/mmroute directory (default: ~/.dsh/mmroute).
  • Model marking synchronization: When marking models on the settings page, the plugin automatically syncs the change into the native input declaration in ~/.dsh/settings.yaml (for example, input: [text, image]), supporting hot route rebuilding.
  • Cache limits: The transcription cache limit is 300 entries, and the marking limit is 2000 entries. Excess entries are evicted using FIFO.
  • Concurrent processing: If multiple concurrent requests transcribe the same image, the calls are merged into one.
  • Understanding model selection: The understanding model can be any provider that declares image input capability. The plugin does not include free endpoints and does not borrow any external login state.

Applicable Scenarios and Notes

This plugin is suitable for scenarios where text-only models (such as DeepSeek or GLM) need to handle visual tasks. By using “transcription” rather than “direct image forwarding,” it ensures that text-only models can understand image content.

Before use, make sure you have correctly configured a multimodal model for “understanding” images (for example, a local Ollama model or another cloud service). The plugin runs with the permissions of the current DSH process. It is recommended to review the source code and license before installation.

Summary

dsh-mmroute bridges the gap between text-only models and visual tasks in the DSH ecosystem. Through an interception layer, it enables integration without changing existing habits, while providing error recovery and structured transcription capabilities. It is currently a relatively professional solution for multimodal routing and agent visual processing.