Introduction¶
When using DeepSeek Harness (DSH), if the session model does not support image input (such as the official DeepSeek model), DSH directly blocks the message and displays “The current model does not support images.” This limits the ability of text-only models to obtain visual information. The dsh-vision-bridge plugin intervenes in DSH’s message processing flow by converting images into tool-call instructions, allowing models that do not support images to analyze images through visual tools.
Plugin Overview¶
This is a DSH profile plugin package maintained by developer ShiraGawaAnri. It aims to address the issue that text-only models such as DeepSeek cannot directly process image input. By invoking visual tools, the plugin bypasses the native blocking mechanism, allowing the model to recognize images and generate responses based on tool results.
Core Features¶
- Image bridge processing: When the model does not support image input, instead of triggering the native blocking behavior, it invokes the configured visual tool (
vision_glanceby default) to recognize the image. - Workspace materialization: Creates a
.dsh-vision-bridge/directory in the session workspace and materializes the image bytes. - Flexible preset control: Supports custom prompts and preset exclusion (such as the default
minimal). Sessions matching an excluded preset retain the native “Image not supported” blocking message. - Cross-platform support: Supports Windows and WSL/Ubuntu platforms and uses the shell to perform image processing.
Installation and Enablement¶
To install the plugin, first extract the downloaded package, then add it to the DSH profile using the following commands.
unzip dsh-vision-bridge.zip -d dsh-vision-bridge
dsh plugin --profile web add ./dsh-vision-bridge
Configuration and Usage¶
After installation, the plugin persists its configuration through the DSH settings service. Find the “Vision Bridge” option in DSH settings to configure it:
- Enable switch: Enabled by default. Disabling it restores DSH’s native blocking behavior.
- Visual tool selection:
vision_glanceis used by default. The system performs fuzzy matching based on the current agent’s toolset, or you can select tools such asglance/ground/detect. - Custom prompt: An optional field used to pass additional instructions to the visual tool.
- Excluded presets: Configure multi-select and wildcard patterns. The default includes the
minimalpreset, and matching sessions will not run the bridge. - Prerequisite: Ensure that the
vision-toolsskill is loaded in the session (or that another visual tool is provided); otherwise, the plugin may degrade or display an informational message.
Notes¶
- Skill dependency: The plugin must rely on the
vision-toolsskill to provide visual tools; otherwise, it cannot perform the bridge properly. - Process-level impact: The plugin modifies
llm.resolveModelInfothrough an Admission patch, which affects all sessions under the current DSH process. Excluded-preset sessions do not run the bridge, but they may not display the native error message and may present it as text instead. - Adapter limitation: The bridge adapter prevents the bridged model from receiving image blocks; images are retained only as attachments in the workspace.
- License: The plugin uses the MIT license. It is recommended to review the source code before installation.
Conclusion¶
By leveraging the tool-call chain, this plugin implements a “text in place of images” bridging logic, giving text models such as DeepSeek visual analysis capabilities. For scenarios where visual information needs to be processed in text models, it is a necessary supplementary tool. For more details, refer to the plugin directory or the GitHub repository.