On October 3, 2026, DeepSeek Harness (DSH) still follows the “everything is a plugin” design philosophy. Text-only models lack the ability to process images, which is a common pain point when building vision agents. The dsh-vision-bridge plugin is designed to add vision capabilities to text-only models. It can either reuse existing DSH Providers or be configured with any OpenAI-compatible multimodal API. This plugin is maintained by alaxrpg.
Core Features¶
The plugin provides the following core capabilities:
vision_bridge_read_imagetool: The model can proactively invoke this tool to identify images. It returns structured textual evidence, including OCR text, layout information, and semantic descriptions.- Automatic pasted-image handling: When an image is pasted into an input control that supports TEXTAREA, INPUT, or contenteditable (DSH 0.1.2-alpha+), the plugin automatically uploads the image and inserts a compact reference marker (for example
「▧ 图片 #KdAy1D」) so the model can later invoke a vision tool. - No vendor hard-coding: Provider and model lists are read in real time from the DSH registry; no vendors are hard-coded.
- Visual configuration: In the “Plugins → Vision Bridge” submenu on the Settings page, you can select a DSH Provider or add a custom OpenAI-compatible Provider. Changes take effect after saving.
- Connectivity test: A one-click test sends a test image to the vision API and returns latency and response samples.
- Reuse existing channels: When a DSH Provider is selected, its adapter, credentials, and attachment channel are reused; no extra wrapper Provider is registered.
Installation and Enabling¶
Use the official CLI to install the plugin (bundles are written automatically):
dsh plugin --profile web add github:alaxrpg/dsh-vision-bridge
After installation, the plugin is loaded into the current DSH process.
Configuration¶
It is recommended to configure it using the GUI settings page. After installing and restarting DSH, open “Settings → Plugins → Vision Bridge.”
Provider Mode¶
- DSH Provider Mode: Select from the Providers and models currently added to DSH. The image is saved in the DSH attachment store and then handed to the selected Provider’s adapter. The plugin does not read the Provider’s Base URL or API key; credentials are managed by the adapter.
- Custom Provider Mode: Configure a direct OpenAI-compatible connection. A custom Provider belongs only to this plugin and is not registered in DSH’s global Providers.
API Key Configuration¶
If using a custom Provider, API keys are resolved with the following priority:
1. vision-bridge.apiKey (entered on the configuration page or supplied directly in settings.yaml).
2. vision-bridge.apiKeyEnv: resolved via the DSH credentials service, taking precedence over environment-variable files.
Important note: The configuration endpoint GET /vision-bridge/config always returns only the boolean status of keySource and keyResolved; it never echoes the key.
Desktop Notes¶
When the desktop host starts, it imports settings.yaml into internal storage once and renames it to settings.yaml.imported. The desktop host does not inject a settings service into this plugin.
To handle this, the plugin provides the following fallback behavior:
* On read, it includes settings.yaml.imported among candidate files.
* If the settings service is unavailable, the “Save” button on the Settings page falls back to writing directly to settings.yaml; changes take effect immediately after the write.
Manual Configuration (YAML)¶
You can also edit ~/.dsh/settings.yaml directly:
vision-bridge:
enabled: true
providerMode: dsh
provider: your-dsh-provider-id
model: your-vision-model-id
timeout: 90
External edits to the configuration file are also watched and applied immediately.
Usage¶
After configuration, the plugin provides two ways to use it:
1. The Model Invokes the Tool¶
The model can proactively invoke the vision_bridge_read_image tool. Example command:
使用 vision_bridge_read_image 工具识别图片:/path/to/image.png
2. Automatic Pasted-Image Handling¶
When a user pastes an image directly, the plugin automatically uploads it and inserts a reference marker. After seeing the marker in context, the model automatically invokes the vision tool to process it.
Output Format¶
The tool call returns structured JSON evidence:
{
"summary": "图片整体描述",
"ocr": {
"full_text": "完整 OCR 文本",
"lines": [{"text": "行文本", "language": "zh"}]
},
"layout": {
"regions": [{"type": "heading", "reading_order": 1, "text": "..."}]
},
"semantics": {
"scene": "场景描述",
"entities": [{"name": "实体名", "type": "类型"}]
},
"uncertainty": ["不确定项"]
}
Technical Details and Compatibility¶
- Compatibility: DSH 0.1.1+, Node.js 18+, macOS / Linux.
- Dependencies:
@deepseek-ai/dsh-credentials(>=0.1.1-rc.1),@deepseek-ai/schemastery(>=3.18.1). - Routes:
GET /vision-bridge/config: Read the configuration status.POST /vision-bridge/config: Save the configuration.GET /vision-bridge/test: Run a connectivity test.POST /vision-bridge/paste: Handle image upload from pasting.GET /vision-bridge/verdict: Determine whether to take over pasting.
- Client: Classic scripts are loaded through
__ModuleLoader__; paste interception continues to work even without a plugin context.