Introduction

DeepSeek Harness (DSH) uses a plugin architecture, and extension behavior is typically implemented by mounting hooks. In Web GUI interaction scenarios, users frequently attach images. If the underlying model is a text-only model (for example, DeepSeek), these images are ignored by default by the system, preventing the model from understanding visual content. The dsh-image-bridge plugin intercepts the hook before message assembly, uses a vision model to convert images into text descriptions, and then injects them into the conversation, giving text-only models the ability to “understand images.”

Plugin Overview

dsh-image-bridge is a DSH plugin maintained by Seryta and licensed under the MIT License. It primarily solves the following problem: when a model does not support image input, it automatically transcribes images from the Web GUI into text using a vision model (by default, Zhipu GLM-4V-Flash) and replaces the original image block before entering the conversation flow.

Core Features

The plugin runs at the agent/pre-step hook. Its core capabilities include:

  1. Interception and Rewriting: Intercepts the agent/pre-step waterfall event, reads image content blocks in the message ({ type: 'image', attachment: ImageAttachmentRef }), and retrieves image bytes via ctx.attachments.readImage.
  2. Automatic Transcription: Converts images into text descriptions and replaces the original image block with an [Image description] text block.
  3. Caching Mechanism: Uses attachmentId as the cache key and records recognized image descriptions. The same image is not re-recognized in subsequent steps (cache limit 100, FIFO eviction). Failed results are not cached.
  4. Resilience Strategy:
    • Per-model retry: On 429/503 errors, retries with exponential backoff (default 2 extra attempts: 1s, 2s).
    • Model fallback: On 404 or exhausted retries (5xx), automatically switches to the next backup model (default chain: glm-4.6v-flash → glm-4.1v-thinking-flash → glm-4v-flash).
    • Failure degradation: If all attempts fail, replaces the image block with [Image recognition failed] text, avoiding UNSUPPORTED_CONTENT failures that would cause the entire conversation turn to fail.
  5. Limits and Protection:
    • Limits image size (default 10MB) and resolution (default 40MP); exceeding the limits causes an immediate error.
    • Handles content parts array robustness; when the gateway returns a [{text}] array, entries are joined in order.
    • Strips the thinking inference block from thinking models.
  6. Recursive Processing: Supports recursive transcription of nested images in tool-result.

Installation and Enablement

Use the following install command:

dsh plugin --profile web add github:Seryta/dsh-image-bridge

After installation, restart the dsh web process to make it take effect.

Configuration

The plugin supports configuration through environment variables, primarily including model selection, retry strategy, rate limit settings, and so on:

Setting Environment Variable Default Value Description
Primary model name IMAGE_BRIDGE_MODEL glm-4v-flash Vision model used for image transcription
Fallback model chain IMAGE_BRIDGE_FALLBACKS glm-4.6v-flash,glm-4.1v-thinking-flash,glm-4v-flash Comma-separated, tried in order
Retry count IMAGE_BRIDGE_RETRIES 2 Extra retries for a single model on 429/503
Description language IMAGE_BRIDGE_LOCALE zh zh Chinese / other English
Image byte limit IMAGE_BRIDGE_MAX_IMAGE_BYTES 10485760 Exceeding the limit fails immediately; no automatic compression
Image pixel limit IMAGE_BRIDGE_MAX_IMAGE_PIXELS 40000000 Exceeding the width×height limit fails immediately

Credential Loading

The plugin attempts to read the vision API key from the following environment variables in order:
* ZHIPU_API_KEY
* ZHIPUAI_API_KEY
* VISION_API_KEY
* DASHSCOPE_API_KEY

Model Limitation Notes

If the selected model declares inputModalities: ['text'], DSH rejects images before messages enter the loop. This plugin itself does not modify model declarations, so it must be paired with a plugin that adds an image modality declaration (such as llm-deepseek-image-admit) to allow all models to accept image input.

Notes

  1. Image processing: Multiple images in a single message are combined into one vision API call. Cache hits are based on ordered attachmentId composite keys; if the image order differs, it is treated as a new request.
  2. No preprocessing: Images are not compressed or pre-scaled; only size and pixel limits are enforced.
  3. Rate limit risks: Free-tier VLMs may have rate limit issues. The plugin provides retry and fallback mechanisms, but API quotas should still be monitored.
  4. Dependency environment: The plugin is handwritten JS with zero build and zero third-party dependencies, but it depends on Node built-in modules.

Summary

dsh-image-bridge provides a bridging solution for visual information from the Web GUI to text-only models. Through interception, transcription, caching, and degradation strategies, it ensures that image information can be understood by models while maintaining conversation stability. It is suitable for developers who need to handle Web GUI image input on text-only models such as DeepSeek.

Plugin Directory
GitHub Repository