Introduction

In DeepSeek Harness (DSH), text-only chat models cannot process images directly. By default, they return a placeholder. nvbb/dsh-ollama-vision-bridge is a plugin that uses a local Ollama vision-language (VL) model (default: qwen3-vl:8b) to describe images and injects the description into the conversation context, allowing the text model to answer based on the image content.

Core Features

  1. Automatic Description: If the selected chat model does not support images, the plugin automatically uses a local Ollama VL model to describe the image.
  2. Synchronous Injection: The description is injected into the model in the same model step.
  3. History Retention: The original image is retained in the conversation history.
  4. VRAM Cooldown: After inference, the model remains in VRAM for a short time and is automatically unloaded when idle.
  5. Patch Management: Applies an idempotent patch to dsh-api-session-controller.
  6. Behavior Fallback: Falls back to DSH’s original rejection behavior.
  7. Status Query: Provides a /vision-bridge status command.
  8. Automatic Mounting: Automatically mounts into the profile’s bundles layer.

Installation and Activation

Before installation, you must stop any running dsh web process; otherwise, node_modules may be locked.

# 从 GitHub 安装
dsh plugin --profile web add git+https://github.com/nvbb/dsh-ollama-vision-bridge.git

# 或者本地开发/调试安装
dsh plugin --profile web add file:<仓库的本地路径>

# 应用补丁(幂等操作,可重复执行)
node "$env:USERPROFILE\.dsh\profiles\node_modules\dsh-ollama-vision-bridge\patch\apply.mjs"

# 更新插件
dsh plugin --profile web update dsh-ollama-vision-bridge

# 检查补丁状态
node patch/apply.mjs --check

# 运行冒烟测试(需本地 Ollama 在运行且模型已下载)
node test/bridge-smoke.mjs

# 确认 Ollama 模型列表
ollama list

Configuration

Configuration items are located in DSH’s settings.yaml file, at the same level as the llm-pi-ai configuration section. The plugin supports hot loading, so no restart is required after configuration changes.

vision-bridge:
  enabled: true                     # 总开关;false 时恢复 DSH 原生拒绝
  baseURL: http://127.0.0.1:11434   # Ollama 服务地址
  model: "qwen3-vl:8b"              # 默认使用的视觉模型
  keepAlive: "60s"                  # 显存冷却时间,空闲后卸载
  # models:                         # 可选:按模型映射不同的视觉模型
  #   "deepseek-v4-flash": "qwen3-vl:8b"

Typical Usage

In the DSH Web GUI, select a text-only model (one that does not support images) and send an image. The plugin automatically recognizes the image content, generates a collapsed “context injection” message in the UI, and then the model provides an answer based on the image description. The original image is retained in the history, but it is not sent directly to the text adapter; instead, the description is obtained via the local VL model.

Applicable Scenarios and Limitations

This plugin is suitable for scenarios that require local deployment, privacy protection, and no dependence on cloud APIs. Before installation, ensure that Ollama is running and contains vision models such as qwen3-vl:8b.

Main limitations are as follows:
1. It depends on two stable anchors in the prompt handler of dsh-api-session-controller; major DSH version changes may cause errors.
2. It only covers the image-sending path for normal chat sessions in the Web GUI; image-rejection behavior for parallel entry points such as subagents is not included.
3. When a VL model (such as qwen3-vl:8b) is selected directly and images are sent, the request goes through DSH’s native channel rather than the bridge logic.

Summary

This plugin solves the problem that text-only models in local DSH environments cannot understand images. By using a local Ollama vision model, it provides local image understanding without relying on public cloud APIs.

  • Directory page: https://www.skillhub.cn/plugins/nvbb/dsh-ollama-vision-bridge
  • GitHub repository: https://github.com/nvbb/dsh-ollama-vision-bridge