Introduction

The DeepSeek v4 model series currently supports text input only. If a user pastes a screenshot or photo into a conversation, the model cannot directly understand the image content. dsh-vision-plugin acts as an intermediate layer: it first sends the image to a vision model (such as qwen3.7-plus, kimi-k3, etc.), transcribes the image into text, and then passes the text to the main model for answering. This allows text models to “see” images and interact naturally.

Identity Information

  • Name: dsh-vision-plugin
  • Core Value: Provides visual analysis capabilities for DeepSeek Harness text models.
  • Owner: Xin-Zhang-IceMan
  • Category: Model inference
  • License: MIT License

Installation

Installing this plugin registers it in the current DSH configuration, so it is loaded automatically when DSH starts.

dsh plugin --profile web add dsh-vision-plugin

Configuration

Before images can be processed, the configuration file must declare that the current model supports image input. The configuration file is located at ~/.dsh/settings.yaml. Changes support hot reloading, so no restart is required.

The following example declares that deepseek-v4-flash and deepseek-v4-pro can accept images:

llm-pi-ai:
  providers:
    opencode-go:
      apiKeyEnv: OPENCODE_GO_API_KEY
      modelOverrides:
        deepseek-v4-flash:
          input: [text, image]
        deepseek-v4-pro:
          input: [text, image]

For DSH v0.1.0-rc.8 and later, the built-in deepseek-official route supports native Vision. Models must declare inputModalities under the llm-deepseek configuration, and the plugin recognizes this automatically.

Core Features

  1. Paste images and ask questions: Directly paste an image into the chat window, switch to any text model, and ask about the image content.
  2. vision_analyze tool: The model can call the tool to directly analyze local image files, with support for an optional question and a specified vision model.
  3. Settings page: A new “Vision Model” page is added to the DSH settings panel for managing the default vision model, with support for switching between Chinese and English.

Typical Usage

Scenario 1: Analyzing images in conversation

Paste an image into the conversation window (for example, the Clash Verge logo), switch to a text model such as deepseek-v4, and enter the question “What is this?”. The plugin automatically calls a vision model to transcribe the image content, and the main model answers based on the transcription.

Scenario 2: Tool-call analysis

Ask the model in the conversation to analyze a table or a code screenshot. The model automatically calls the vision_analyze tool. The tool accepts a local image path, an optional question, and an optional vision model override parameter.

Notes

  • Version requirement: The plugin is compatible with dsh v0.1.0-rc.8 and later, and supports native Vision routes.
  • Configuration validity: If you switch to a model that has not declared image support in settings.yaml, pasting an image may be rejected. Make sure all text models that may be used have declared input: [text, image].
  • State persistence: The vision model selection on the settings page is stored in the browser’s localStorage and is restored automatically after DSH restarts.
  • Language sync: The UI language of the settings page follows the global DSH language setting (Chinese or English).
  • GitHub repository: https://github.com/Xin-Zhang-IceMan/dsh-vision-plugin
  • Plugin directory: https://www.skillhub.cn/plugins/Xin-Zhang-IceMan/dsh-vision-plugin