Introduction¶
DeepSeek Harness (DSH) uses a plugin architecture, and extension behavior is typically implemented by mounting hooks. In Web GUI interaction scenarios, users frequently attach images. If the underlying model is a text-only model (for example, DeepSeek), these images are ignored by default by the system, preventing the model from understanding visual content. The dsh-image-bridge plugin intercepts the hook before message assembly, uses a vision model to convert images into text descriptions, and then injects them into the conversation, giving text-only models the ability to “understand images.”
Plugin Overview¶
dsh-image-bridge is a DSH plugin maintained by Seryta and licensed under the MIT License. It primarily solves the following problem: when a model does not support image input, it automatically transcribes images from the Web GUI into text using a vision model (by default, Zhipu GLM-4V-Flash) and replaces the original image block before entering the conversation flow.
Core Features¶
The plugin runs at the agent/pre-step hook. Its core capabilities include:
- Interception and Rewriting: Intercepts the
agent/pre-stepwaterfall event, reads image content blocks in the message ({ type: 'image', attachment: ImageAttachmentRef }), and retrieves image bytes viactx.attachments.readImage. - Automatic Transcription: Converts images into text descriptions and replaces the original image block with an
[Image description]text block. - Caching Mechanism: Uses
attachmentIdas the cache key and records recognized image descriptions. The same image is not re-recognized in subsequent steps (cache limit 100, FIFO eviction). Failed results are not cached. - Resilience Strategy:
- Per-model retry: On 429/503 errors, retries with exponential backoff (default 2 extra attempts: 1s, 2s).
- Model fallback: On 404 or exhausted retries (5xx), automatically switches to the next backup model (default chain:
glm-4.6v-flash → glm-4.1v-thinking-flash → glm-4v-flash). - Failure degradation: If all attempts fail, replaces the image block with
[Image recognition failed]text, avoidingUNSUPPORTED_CONTENTfailures that would cause the entire conversation turn to fail.
- Limits and Protection:
- Limits image size (default 10MB) and resolution (default 40MP); exceeding the limits causes an immediate error.
- Handles
content partsarray robustness; when the gateway returns a[{text}]array, entries are joined in order. - Strips the
thinkinginference block from thinking models.
- Recursive Processing: Supports recursive transcription of nested images in
tool-result.
Installation and Enablement¶
Use the following install command:
dsh plugin --profile web add github:Seryta/dsh-image-bridge
After installation, restart the dsh web process to make it take effect.
Configuration¶
The plugin supports configuration through environment variables, primarily including model selection, retry strategy, rate limit settings, and so on:
| Setting | Environment Variable | Default Value | Description |
|---|---|---|---|
| Primary model name | IMAGE_BRIDGE_MODEL |
glm-4v-flash |
Vision model used for image transcription |
| Fallback model chain | IMAGE_BRIDGE_FALLBACKS |
glm-4.6v-flash,glm-4.1v-thinking-flash,glm-4v-flash |
Comma-separated, tried in order |
| Retry count | IMAGE_BRIDGE_RETRIES |
2 |
Extra retries for a single model on 429/503 |
| Description language | IMAGE_BRIDGE_LOCALE |
zh |
zh Chinese / other English |
| Image byte limit | IMAGE_BRIDGE_MAX_IMAGE_BYTES |
10485760 |
Exceeding the limit fails immediately; no automatic compression |
| Image pixel limit | IMAGE_BRIDGE_MAX_IMAGE_PIXELS |
40000000 |
Exceeding the width×height limit fails immediately |
Credential Loading¶
The plugin attempts to read the vision API key from the following environment variables in order:
* ZHIPU_API_KEY
* ZHIPUAI_API_KEY
* VISION_API_KEY
* DASHSCOPE_API_KEY
Model Limitation Notes¶
If the selected model declares inputModalities: ['text'], DSH rejects images before messages enter the loop. This plugin itself does not modify model declarations, so it must be paired with a plugin that adds an image modality declaration (such as llm-deepseek-image-admit) to allow all models to accept image input.
Notes¶
- Image processing: Multiple images in a single message are combined into one vision API call. Cache hits are based on ordered
attachmentIdcomposite keys; if the image order differs, it is treated as a new request. - No preprocessing: Images are not compressed or pre-scaled; only size and pixel limits are enforced.
- Rate limit risks: Free-tier VLMs may have rate limit issues. The plugin provides retry and fallback mechanisms, but API quotas should still be monitored.
- Dependency environment: The plugin is handwritten JS with zero build and zero third-party dependencies, but it depends on Node built-in modules.
Summary¶
dsh-image-bridge provides a bridging solution for visual information from the Web GUI to text-only models. Through interception, transcription, caching, and degradation strategies, it ensures that image information can be understood by models while maintaining conversation stability. It is suitable for developers who need to handle Web GUI image input on text-only models such as DeepSeek.