Preface

DeepSeek Harness (DSH) uses a host-level plugin architecture to extend functionality. Pure text models (such as DeepSeek V4) cannot process images directly and are prone to making assumptions about image contents out of thin air. dsh-visibridge is a host-level vision plugin that provides two tools: analyze_image and capture_image. It sends images to a configured vision model (local Ollama or a cloud API) and returns structured JSON evidence (including full OCR text, layout analysis, entity recognition, etc.), allowing the pure text model to reason based on this evidence rather than guessing.

Core Features

The plugin primarily bridges the gap between pure text models and vision capabilities, providing the following features:

  • Image recognition: Supports local file paths and HTTP(S) URLs.
  • Structured evidence: Returns structured data such as summary (overview), ocr (full text + line-by-line), layout (layout), and semantics (scene + entities).
  • Multi-backend support: Built-in preset backends, switchable with one line.
    • ollama: Local processing; images do not leave the local machine.
    • xiaomi: Xiaomi MiMo cloud.
    • deepseek: DeepSeek official vision model.
    • custom: Any OpenAI-compatible endpoint.
  • Key security: API keys in error logs are automatically redacted to [REDACTED].

Installation

The plugin is published on npm. It is recommended to install it using the official command.

dsh plugin --profile web add dsh-visibridge

This command automatically loads the plugin into the bundles list of the profile. It takes effect after restarting dsh.

Configuration

The plugin supports specifying the backend and model through configuration files. Configuration priority is: plugin defaults → cordis.yml configuration → workspace dsh-vision-config.json.

1. Workspace Configuration File

Create dsh-vision-config.json in the workspace root:

{
  "backend": "ollama",
  "model": "minicpm-v4.5",
  "timeoutMs": 300000
}

2. Backend Presets

The backend field determines the default configuration:

  • ollama: Local endpoint; the recommended model is minicpm-v4.5.
  • xiaomi: Cloud endpoint; an API key must be configured.
  • deepseek: Official vision model.

3. Credential Configuration

Configure the cloud API key in ~/.dsh/.credentials.yaml:

XIAOMI_MIMO_API_KEY: sk-xxxxxxxx

Usage

After installing and restarting dsh, the tools are automatically registered.

Image Analysis

  1. Place the image in the session workspace.
  2. Send instructions to the AI, for example: look at this image or analyze xxx.png.
  3. The AI invokes the analyze_image tool and answers based on the structured evidence.

Camera Capture and Recognition

After connecting a USB camera, the AI can autonomously invoke the capture_image tool. For example, tell the AI:
* “Take a photo and check the phone screen.”
* “See what error is currently displayed on the screen.”

The AI will automatically capture, recognize, and return structured evidence. The captured images are saved in the workspace .captures/ directory.

Hardware and Notes

  • Environment requirements:
    • Node.js >= 22.
    • If using Ollama, install the corresponding vision model.
  • Camera requirements (used only by capture_image):
    • USB interface (UVC driverless).
    • Supports autofocus.
    • Minimum focusing distance ≤ 10 cm (critical for capturing phone screens/documents).
  • Security policies:
    • baseUrl rejects internal addresses by default (RFC1918, link-local, etc.) to prevent data leakage. To connect to an internal endpoint, explicitly enable "allowPrivateHosts": true in the configuration.
    • The workspace dsh-vision-config.json can be modified by sessions. Use it only in trusted directories.

Conclusion

dsh-visibridge bridges pure text models and vision capabilities through structured evidence. For agent development that requires document processing, OCR, or visual feedback, it provides a localized and configurable solution.