In DeepSeek Harness (DSH), when interacting with text-only chat models, users cannot natively attach images. Starting from DSH 0.1.2-alpha.2+, the core session controller enforces strict modality checks and directly rejects attachments if the current model does not support image input. dsh-vision-bridge is a runtime intermediary plugin that enables text models to process visual content by intercepting and rewriting image blocks.

Core Features

The plugin provides the following features, enabling visual capabilities with text-only models:

  • Server-side modality bridging (v0.5.3+): By decorating ctx.llm.resolveModelInfo and ctx.llm.listModels, it allows the session controller to accept image attachments for all models.
  • Automatic image rewriting: It intercepts image blocks via agent/pre-step and llm/stream, routes them to the configured vision model, receives descriptive synthetic text, and rewrites it as text context (e.g., [The user attached an image. Description: ...]) before passing it to the text model.
  • Native passthrough: It automatically detects models that natively support vision and allows images to pass through directly without rewriting.
  • Rich visual toolset: It exposes about 40 dedicated tools, including OCR, visual question answering (VQA), grounding, UI flows, consensus, and more.
  • Multi-channel endpoint routing: It supports dsh-catalog, OpenAI-compatible, Ollama, and webhook endpoints, with failover.
  • High-performance LRU description cache: It caches vision responses by hash to eliminate redundant API calls and conserve quota.

Installation

Run the following command in the terminal to install the plugin:

dsh plugin --profile web add @goodandready/dsh-vision-bridge

After installation, you must restart the DSH Web UI. The configuration card can be found under Settings -> Plugins -> vision-bridge.

Configuration and Usage

The plugin supports three operating modes:

  1. hybrid (default): Automatically describes attached images during chat turns while retaining all ~40 explicit visual tools for subsequent inference.
  2. llm: Pure automatic rewriting mode; images are transparently converted into text context, and tools remain callable.
  3. tools: Disables automatic rewriting; the chat model must explicitly call describe_image or an OCR tool.

Configuration path: Settings → Plugins → vision-bridge.

Toolset Overview

The plugin provides about 40 vision-related tools covering the following categories:

  • Core: describe_image, read_image, inspect_image
  • Geometry and detection: vision_ground, vision_crop, vision_detect, etc.
  • OCR and text: vision_ocr, vision_ocr_local, vision_long_ocr, etc.
  • Structured and UI: vision_describe_structured, vision_vqa, vision_ui_layout, etc.
  • Documents and intelligence: vision_extract_formula, vision_extract_table, vision_scan_barcode, etc.
  • Scenes and consensus: vision_ui_flow, vision_consensus, vision_memory_search, etc.

Ecosystem Context

The core philosophy of DeepSeek Harness is “everything is a plugin”. The plugin maintainer is listed in the community catalog (skillhub.cn) and has no official affiliation with DeepSeek.

Notes:
* A Node.js 20+ environment is required.
* Be sure to restart the Web UI after installation.
* The plugin is open source under the MIT License.

Project URL: GitHub