Preface¶
DeepSeek Harness (DSH) uses a host-level plugin architecture to extend functionality. Pure text models (such as DeepSeek V4) cannot process images directly and are prone to making assumptions about image contents out of thin air. dsh-visibridge is a host-level vision plugin that provides two tools: analyze_image and capture_image. It sends images to a configured vision model (local Ollama or a cloud API) and returns structured JSON evidence (including full OCR text, layout analysis, entity recognition, etc.), allowing the pure text model to reason based on this evidence rather than guessing.
Core Features¶
The plugin primarily bridges the gap between pure text models and vision capabilities, providing the following features:
- Image recognition: Supports local file paths and HTTP(S) URLs.
- Structured evidence: Returns structured data such as
summary(overview),ocr(full text + line-by-line),layout(layout), andsemantics(scene + entities). - Multi-backend support: Built-in preset backends, switchable with one line.
ollama: Local processing; images do not leave the local machine.xiaomi: Xiaomi MiMo cloud.deepseek: DeepSeek official vision model.custom: Any OpenAI-compatible endpoint.
- Key security: API keys in error logs are automatically redacted to
[REDACTED].
Installation¶
The plugin is published on npm. It is recommended to install it using the official command.
dsh plugin --profile web add dsh-visibridge
This command automatically loads the plugin into the bundles list of the profile. It takes effect after restarting dsh.
Configuration¶
The plugin supports specifying the backend and model through configuration files. Configuration priority is: plugin defaults → cordis.yml configuration → workspace dsh-vision-config.json.
1. Workspace Configuration File¶
Create dsh-vision-config.json in the workspace root:
{
"backend": "ollama",
"model": "minicpm-v4.5",
"timeoutMs": 300000
}
2. Backend Presets¶
The backend field determines the default configuration:
ollama: Local endpoint; the recommended model isminicpm-v4.5.xiaomi: Cloud endpoint; an API key must be configured.deepseek: Official vision model.
3. Credential Configuration¶
Configure the cloud API key in ~/.dsh/.credentials.yaml:
XIAOMI_MIMO_API_KEY: sk-xxxxxxxx
Usage¶
After installing and restarting dsh, the tools are automatically registered.
Image Analysis¶
- Place the image in the session workspace.
- Send instructions to the AI, for example:
look at this imageoranalyze xxx.png. - The AI invokes the
analyze_imagetool and answers based on the structured evidence.
Camera Capture and Recognition¶
After connecting a USB camera, the AI can autonomously invoke the capture_image tool. For example, tell the AI:
* “Take a photo and check the phone screen.”
* “See what error is currently displayed on the screen.”
The AI will automatically capture, recognize, and return structured evidence. The captured images are saved in the workspace .captures/ directory.
Hardware and Notes¶
- Environment requirements:
- Node.js >= 22.
- If using Ollama, install the corresponding vision model.
- Camera requirements (used only by
capture_image):- USB interface (UVC driverless).
- Supports autofocus.
- Minimum focusing distance ≤ 10 cm (critical for capturing phone screens/documents).
- Security policies:
baseUrlrejects internal addresses by default (RFC1918, link-local, etc.) to prevent data leakage. To connect to an internal endpoint, explicitly enable"allowPrivateHosts": truein the configuration.- The workspace
dsh-vision-config.jsoncan be modified by sessions. Use it only in trusted directories.
Conclusion¶
dsh-visibridge bridges pure text models and vision capabilities through structured evidence. For agent development that requires document processing, OCR, or visual feedback, it provides a localized and configurable solution.
- Project directory: dsh-visibridge
- Source code: GitHub