Introduction¶
When developing agents based on DeepSeek Harness (DSH), we often encounter a specific problem: the project is configured with text-only models such as DeepSeek or GLM, but when handling requests that include images (such as screenshot analysis or chart reading), image rejection or errors occur. Manually switching routes can solve the problem, but it disrupts the agent’s continuity and established behavior.
As a DSH plugin, the core value of dsh-mmroute lies in solving this problem “transparently.” It does not require changes to existing model invocation habits. Instead, it intervenes in the llm/stream flow as an intermediate layer, intercepting, transcribing, and dispatching images in requests sent by text-only models, ensuring that the entire agent flow can handle visual tasks smoothly.
Plugin Positioning¶
- Name: dsh-mmroute
- Maintainer: jmxsxwyzjdwl
- Category: Model inference
- License: MIT
- Core capability: Provides image modality scheduling for every model route in DSH, across the entire agent flow.
Core Capabilities¶
The plugin mainly handles the following scenarios and functions:
-
End-to-end image modality scheduling
From user-uploaded images andread_imageresults to rendered images produced by MCP tools (such as Figma screenshots), all images are handled uniformly. -
Automatic transcription for text-only models
For text-only models (or models not marked as multimodal), the plugin transcribes images in their requests into detailed textual descriptions (including verbatim transcription, chart data, layout, and color information). -
Automatic recovery from image-related errors
If a text-only model (or gateway) rejects images and causes the request to fail, the plugin automatically reroutes to an understanding model and retries until success or the retry limit is reached. -
Historical image replay with summaries
When the same image appears again in the conversation history, it is replayed as a summary (≤1200 characters), preventing long sessions from bloating due to full resending of old images. -
Structured transcription and
vision_relook
Transcriptions include structured information such as image type classification, full text content, and chart data. When responding, text-only models can invoke thevision_relooktool to re-examine specific image details, forming a “command-and-execution” collaboration loop. -
Multi-format support
Supports common image formats such as PNG, JPEG, WebP, GIF.
Installation and Enablement¶
The installation process is straightforward and can be completed using the official command:
dsh plugin --profile web add dsh-mmroute
After installation, restart the DSH Web process for the plugin to take effect. The plugin automatically writes dependencies into the profile’s package.json.
Usage¶
After enabling the plugin, the workflow is as follows:
- Open Settings → Multimodal Routing.
- In the “Multimodal Understanding Model” dropdown, select:
- Automatic: Use the first discovered native multimodal model.
- Specify: Manually select a configured multimodal model (it can be a free cloud model or a local model).
- (Optional) In “Model Modality Marking”, explicitly select Default / Multimodal / Text-only for individual models. This overrides the automatic detection result.
- In the conversation, directly paste or upload an image, or let the agent invoke MCP tools to generate images, and you will see the automatic transcription effect.
Note: Starting with v0.6.0, the settings page has been simplified and only retains the understanding model and modality marking. Features such as automatic routing, error recovery, and historical image summaries are enabled by default and run continuously in the background.
Configuration and Mechanism Details¶
- Configuration persistence: The plugin’s configuration is stored in the
$DSH_HOME/mmroutedirectory (default:~/.dsh/mmroute). - Model marking synchronization: When marking models on the settings page, the plugin automatically syncs the change into the native
inputdeclaration in~/.dsh/settings.yaml(for example,input: [text, image]), supporting hot route rebuilding. - Cache limits: The transcription cache limit is 300 entries, and the marking limit is 2000 entries. Excess entries are evicted using FIFO.
- Concurrent processing: If multiple concurrent requests transcribe the same image, the calls are merged into one.
- Understanding model selection: The understanding model can be any provider that declares image input capability. The plugin does not include free endpoints and does not borrow any external login state.
Applicable Scenarios and Notes¶
This plugin is suitable for scenarios where text-only models (such as DeepSeek or GLM) need to handle visual tasks. By using “transcription” rather than “direct image forwarding,” it ensures that text-only models can understand image content.
Before use, make sure you have correctly configured a multimodal model for “understanding” images (for example, a local Ollama model or another cloud service). The plugin runs with the permissions of the current DSH process. It is recommended to review the source code and license before installation.
Summary¶
dsh-mmroute bridges the gap between text-only models and visual tasks in the DSH ecosystem. Through an interception layer, it enables integration without changing existing habits, while providing error recovery and structured transcription capabilities. It is currently a relatively professional solution for multimodal routing and agent visual processing.