Introduction¶
In DeepSeek Harness, directly pasting an image into a text-only model is rejected, with an error indicating that image input is not supported. The @zzdream67/dsh-vision-bridge plugin intercepts the llm/stream pipeline and replaces image blocks with transcription text from a vision model before the request reaches the upstream provider. This allows a pure-text model to receive text containing image descriptions and still complete inference tasks.
Installation and Activation¶
Installing the plugin requires ensuring that the dsh command and pnpm are installed in the environment (dsh plugin depends internally on pnpm). It is recommended to use npm to install the pre-built version.
- Install
pnpm(if it is not already installed):
npm install -g pnpm
*Note: `corepack enable pnpm` may sometimes fail on Windows due to write permission issues. Installing globally via npm is more reliable.*
- Run the installation command:
dsh plugin --profile web add @zzdream67/dsh-vision-bridge
Core Features¶
The plugin mainly provides the following capabilities:
* Image interception and replacement: It intercepts the LLM pipeline and replaces image content with transcription from a vision model.
* Model declaration injection: It automatically adds an input: [text, image] declaration to bridged models so they can pass DeepSeek Harness admission checks.
* Web settings page: It registers a custom section in the Web settings page to provide a visual configuration interface.
Configuration and Usage¶
The plugin registers a custom section in settings.section of the Web settings page. On that page, you can manage the vision model through three buttons:
- Test and Enable: Sends a 16x16 PNG image to the selected vision model. The declaration is written only when the model accepts the image; if the test fails, it is automatically revoked.
- Enable Directly: Skips the test step and writes the declaration directly. Suitable for scenarios where the model’s vision capability has already been confirmed.
- Reload: Reloads the configuration from disk. Because the desktop client has no reload gesture, this is the only way to update the configuration.
Field Description¶
In the configuration section, you can adjust the following parameters:
| Field | Default Value | Description |
|---|---|---|
enabled |
true |
Master switch. When disabled, the plugin remains loaded but no longer intervenes and revokes the declaration. |
visionProvider |
'' |
Vision route ID. If left empty, the plugin is ineffective. |
visionModel |
'' |
Model ID in the vision route. If left empty, the plugin is ineffective. |
prompt |
- | The prompt sent to the vision model. |
bridge |
[] |
List of pure-text model IDs to be bridged. |
manageDeclarations |
true |
Whether the plugin automatically writes/revokes modality declarations. |
cacheSize |
64 |
Number of cached transcription results. Set to 0 to disable caching. |
timeoutMs |
120000 |
Timeout for transcribing a single image. |
Declaration Mechanism¶
The DeepSeek Harness admission gate (dsh-host-apiproxy) checks a model’s inputModalities. If a model does not declare image support, the request is rejected and the plugin cannot intervene.
- When
manageDeclarationsis enabled, the plugin automatically adds aninput: [text, image]declaration to the models in thebridgelist. - When the plugin is enabled and a model is bridged, that model can effectively receive image input, because the transformation occurs before the provider is called.
- A declaration is an “admission tag”. It only affects whether the host layer allows images into a session, and does not affect Token counting, context windows, or sampling behavior.
Notes and Implementation Details¶
- Form rendering: Registering the schema on the host side alone does not automatically render the form. The plugin implements the UI by manually writing a browser-side bundle.
- Uninstall behavior:
- The plugin records the input state before declarations. If it detects that the input has been manually modified to an incorrect format during uninstallation, it preserves that modification.
- If a model or route has been deleted, the plugin skips that entry (it does not recreate it from nothing).
- The plugin attempts to restore the declaration state for all other routes.
- Source code and license: The plugin uses the MIT license. It is recommended to inspect the source code before installation to ensure it meets security requirements.
Summary¶
This plugin solves the pain point that pure-text models in DeepSeek Harness cannot handle image input. By inserting a vision transcription layer into the LLM pipeline, it enables text models to indirectly understand image content. The related documentation and source code can be viewed on GitHub.