Preface¶
The design philosophy of DeepSeek Harness (DSH) is “everything is a plugin.” When the active text model cannot directly process image inputs, the zh-u-hb/dsh-vision-bridge plugin can step in. It listens to DSH’s message-processing pipeline, and when a text model reports that image input is not supported, it forwards image messages to a user-configured OpenAI-compatible vision model and replaces the returned text back into the conversation flow.
Plugin Overview¶
This is a network-enabled utility plugin maintained by Zh-U-hB. It provides a web settings page that allows users to configure an OpenAI-compatible API endpoint and key. When activated, the plugin intercepts image messages, sends them to the configured vision model, and replaces the returned text back into the conversation flow.
Core Features¶
- Hosting and Configuration: Provides a host plugin and a web settings page.
- Message Routing: Listens to the
agent/pre-steppipeline and forwards image messages to an external endpoint when the active text model reports that image input is not supported. - Prompt Composition: Automatically asks the text model to write a precise prompt for the image, then sends the system prompt combined with the image data.
Installation and Enablement¶
The installation script installs the plugin into the default web configuration file. A service restart is required after installation.
- Run the installation command:
./install.sh
To install into another configuration profile, you can use an environment variable:
DSH_PROFILE=myprofile ./install.sh
- Restart the DSH Web service:
dsh web
Configuration Steps¶
Access the Vision section under the vision-bridge namespace through the web settings page.
The following fields must be configured to enable the bridge:
* enabled: Whether to enable the image bridging feature.
* url: The base address of the OpenAI-compatible API (or the full address, which must end with /chat/completions).
* apiKey: The Bearer Token used for authentication.
* model: The ID of the vision model to invoke.
The settings page masks the API Key. Empty fields are treated as “keep the saved key.”
Working Principle¶
The plugin listens to the agent/pre-step pipeline. Before a message is appended to the session log, the system checks whether the message contains image content blocks.
- If the currently active model explicitly supports image input, the message passes through unchanged.
- If the model only supports text and an endpoint is configured, the plugin performs the following actions:
* Ask the current text model to create a precise visual prompt based on the message text.
* Assemble the system prompt.
* Send a POST request to the configured OpenAI-compatible endpoint, including the system prompt, the generated prompt, and the Base64-encoded image data.
* Replace the original image block with the returned text.
Notes and Limitations¶
- OpenAI Protocol Only: Only the standard
/chat/completionsendpoint is supported. - Image Source Limitation: Only image messages received during the
agent/pre-stepstage are processed. If an image has already reached the session log through another path (for example, tool execution results), the plugin will not rewrite it. - Prerequisite: The active model must be explicitly declared as
text-onlyin its metadata to trigger bridging. If a model does not declare modality information, the message remains unchanged. - Key Management: The web interface cannot directly clear a saved API Key. You must edit the raw configuration document directly.
Summary¶
This plugin injects logic into DSH’s pipeline to solve the problem of text models being unable to process images directly. After configuration, it can seamlessly delegate image-analysis tasks to an external vision model while preserving the reconstructability of the conversation record.