Introduction

In DeepSeek Harness (DSH), many text models (such as DeepSeek) have inputModalities limited to text only. When you try to paste an image in a conversation, the DSH api-proxy directly rejects image-containing messages at the sending stage (with an error such as “the current model does not support images”). The dsh-image-bridge plugin removes this limitation, enabling text models to receive and process images in the DSH environment.

Plugin Overview

Role: A DSH plugin maintained by developer haitang1.
Core value: Allows text models to correctly paste and send images in DSH.

Core Features

  1. Image persistence: Images are automatically written to the .attachments/ directory in the session workspace.
  2. UI display: The conversation area shows an image thumbnail (clickable to enlarge), without displaying the file path text.
  3. Model input: The model receives text of the form [ImageN]:"<path>", allowing it to invoke existing vision or MCP tools (such as vision_glance or mcp__mcp-vision__analyze_image).
  4. Native support: Models that natively support images are unchanged and continue using the native path.
  5. Technical implementation:
    - Wraps llm.resolveModelInfo to allow the precheck to pass.
    - Uses attachments.readImage in agent/pre-step to persist the image.
    - Uses agent/request-error to decouple display from model input.

Installation and Activation

This package is a profile bundle (package.json declares dsh.bundle.patch). Run the following command from the root of a DSH source checkout:

pnpm dsh plugin --profile web add 'github:haitang1/dsh-image-bridge#b5c64ea'

After successful installation, the plugin is automatically added to the profile’s dsh.profile.bundles and takes effect on the next startup (or after a hot load of a long-running surface).

If the CLI is unavailable, you can perform the equivalent steps manually:
1. Edit $DSH_HOME/profiles/<profile>/package.json:
- Add to dependencies: "dsh-image-bridge": "github:haitang1/dsh-image-bridge#b5c64ea"
- Add to the end of the dsh.profile.bundles array: "dsh-image-bridge"
2. Run pnpm install in that profile directory.

Principles and Implementation

Text models are usually rejected by the api-proxy. This plugin enables support through three steps:

  1. Allow the precheck: Wraps llm.resolveModelInfo, adds the image modality for text models, and allows image-containing messages to enter the agent flow.
  2. Persist the image: In agent/pre-step, uses attachments.readImage to read the image bytes and writes them to .attachments/<sha256>.<ext> in the workspace.
  3. Decouple display from model input: Uses DSH’s surface replacement mechanism to append a “model-visible only” replacement event in agent/request-error, replacing the image block in the model-visible surface with [ImageN]:"<path>", while the human transcript still renders the original image. The first call fails locally because the adapter rejects images (with no API cost), after which it automatically retries.

Usage Notes

  1. Limitation on mixing commands and images: Do not mix commands (starting with /) and images in the same turn. When the DSH client input box submits a “command path” (such as /plan or /goal), it does not include draft images; the image remains in the input box. Workarounds: send the command first, then send the image separately; or send the command and image as two separate messages.
  2. Handling historical images: Image blocks left in the conversation history that were generated by vision models (before switching models) are not handled by this plugin when you switch back to a text model, and may trigger adapter errors. It is recommended to clear or compress the history before switching back to a text model.

Summary

By wrapping interfaces and leveraging DSH’s surface replacement mechanism, the plugin enables pure text models to process images through MCP tools. For development scenarios involving text models that need visual capabilities, this is an essential utility. For more details, visit GitHub.