The core design philosophy of DeepSeek Harness is “everything is a plugin”. When building agents or workflows, you often encounter scenarios where a pure text routing model (such as DeepSeek) needs to handle image attachments. Sending images directly may cause errors or make the model unable to understand them, while forcing conversion to Base64 increases token consumption. The dsh-vision plugin addresses this pain point by adding a bridge layer that enables pure text models to “see” and understand images.

Plugin Introduction

dsh-vision is a Vision Bridge plugin for DeepSeek Harness, maintained by user lakeofsky347. It sends user-uploaded images to a dedicated vision model to generate a text description, then inserts the description back for the pure text model, avoiding whole-turn conversation failures caused by unsupported content.

Core Features

  • Image-to-text conversion: Automatically sends image blocks to a vision model (default xiaomi/mimo-v2-omni) to generate descriptions, replacing the original image blocks.
  • Memoized caching: Caches descriptions by (session, attachment ID), reuses text during history replay, and avoids repeated billing.
  • Fault-tolerant fallback: If the vision model call fails, it falls back to placeholder text to ensure the DeepSeek turn does not break.
  • Concurrency isolation: Uses WeakSet to mark different turns; no global state, so they do not interfere with each other.
  • Bundle mode: After installation, the plugin entry is automatically inserted; no manual configuration is required.
  • Zero-overhead passthrough: If the request contains no image blocks, or the routing model natively supports images, it is passed through unchanged.

Installation and Activation

The plugin is published as a Bundle. The installation command is:

dsh plugin --profile web add github:lakeofsky347/dsh-vision

Notes

  • Environment requirements: pnpm version ≥ 10 is required.
  • Build permissions: When installing from GitHub for the first time, the prepare script for the git dependency is disabled by default, causing installation to fail. In this case, add allowBuilds.dsh-vision: true to pnpm-workspace.yaml in the profile, then run the installation command again.
  • License: MIT License.

Configuration and Usage

After installation, the plugin automatically inserts the configuration entry. You need to override the configuration in your local ~/.dsh/profiles/<name>/cordis.patch.yml file to specify the vision model’s provider and model:

- id: vision
  name: dsh-vision
  config:
    provider: xiaomi-token-plan-cn
    model: mimo-v2.5

The configuration items are described as follows:

Configuration Item Default Value Description
provider xiaomi The routing provider where the vision model is located
model mimo-v2-omni Vision model ID
prompt (see source code) The instruction sent with the image to the vision model
cache true Whether to memoize/cache descriptions by session and attachment
maxTokens - Output token limit for the vision call (optional)

When handling image attachments directly in a pure text model session, the plugin automatically intervenes and completes the conversion.

Technical Principles

The plugin registers a listener on the llm/stream waterfall and adopts different strategies based on the situation:

  1. No image blocks: Calls next() directly to pass the request through unchanged (zero overhead).
  2. Native image support: If the routing model declares the image input modality (for example, routing directly to mimo), the request is allowed through as-is.
  3. Pure text model + image: The plugin intercepts the chain and returns a lazy async generator:
    • Parses the routing model’s input modalities.
    • Iterates through the messages and calls the vision model once for content containing image blocks (cached based on (session, attachment ID)).
    • Replaces the image blocks with a text block such as [a visual description of the user's attached image], then reconstructs the request.
    • Re-enters ctx.llm.stream() with a reentry marker to prevent triggering the listener again.

Applicable Scenarios and Precautions

This plugin is suitable for scenarios that require handling image attachments in a pure text routing model. Because the plugin runs with the permissions of the current process, be sure to check the source code and license before installation. For more details, refer to the plugin catalog page or the GitHub repository.

Plugin Catalog | GitHub Repository