Introduction

The philosophy of DeepSeek Harness is “everything is a plugin.” When building agents, pure text models cannot directly handle image inputs. This plugin solves this problem: it intercepts messages containing images and determines whether the current model natively supports image. Multimodal models are passed through directly, while pure text models call a visual model to generate a description and then replace image blocks with text. This way, pure text models can also “understand” images.

Installation and Enablement

This plugin is deployed to the Harness home directory as a single-file ESM plugin and enabled through a home-level patch. It applies to all profiles (web / headless, and any newly created profiles in the future).

Windows (PowerShell)

# 1. 部署插件代码
$dest = "$env:USERPROFILE\.dsh\profiles\node_modules\deepseek-visual-plugin"
New-Item -ItemType Directory -Path $dest -Force | Out-Null
Copy-Item index.js, package.json, cordis.patch.yml -Destination $dest -Force

# 2. 在 $DSH_HOME/cordis.patch.yml 中加入(不存在则创建)
# - insert:
#     - id: deepseek-visual
#       name: deepseek-visual-plugin
#       config:
#         model: <视觉模型名>
#         baseUrl: <OpenAI 兼容端点>
#         apiKeyEnv: <凭证引用名>

Linux / macOS (Bash / Zsh)

# 1. 部署插件代码
dest="$HOME/.dsh/profiles/node_modules/deepseek-visual-plugin"
mkdir -p "$dest"
cp index.js package.json cordis.patch.yml "$dest"

# 2. 在 $DSH_HOME/cordis.patch.yml 中加入
# - insert:
#     - id: deepseek-visual
#       name: deepseek-visual-plugin
#       config:
#         model: <视觉模型名>
#         baseUrl: <OpenAI 兼容端点>
#         apiKeyEnv: <凭证引用名>

It can also be installed as a regular bundle into a single profile:

dsh plugin --profile web add C:\path\to\deepseek-visual-plugin

Configuration

The plugin is configured through cordis.patch.yml. The adjustable parameters are as follows:

Field Default Description
model qwen3.7-plus Visual model name (model ID on an OpenAI-compatible endpoint)
baseUrl DashScope-compatible endpoint When left empty, falls back to QWEN_BASE_URL / DASHSCOPE_BASE_URL / OPENAI_BASE_URL
apiKey Empty Explicit API key override (optional; writing it in plaintext configuration is not recommended)
apiKeyEnv Empty Credential reference name (e.g., DOUBAO_API_KEY), resolved through dsh’s credentials service
imagePrompt Built-in Chinese prompt Image description prompt
timeoutMs 60000 Timeout for a single visual request

API Key resolution priority: config.apiKey (explicit override) → config.apiKeyEnv (credentials service resolving ~/.dsh/.credentials.yaml) → environment variable fallback (QWEN_API_KEY / DASHSCOPE_API_KEY / OPENAI_API_KEY).

Recommended practice: store the key in ~/.dsh/.credentials.yaml and write only the apiKeyEnv reference in the configuration, avoiding plaintext keys in the configuration file.

# ~/.dsh/.credentials.yaml
DOUBAO_API_KEY: <你的 key>
# cordis.patch.yml
- insert:
    - id: deepseek-visual
      name: deepseek-visual-plugin
      config:
        model: doubao-seed-2-0-mini-260428
        baseUrl: https://ark.cn-beijing.volces.com/api/v3
        apiKeyEnv: DOUBAO_API_KEY

How It Works

The plugin is implemented against the native protocol and includes the following main logic:

  1. Interception and detection: Intercepts messages containing images and checks whether the current model natively supports image.
  2. Routing:
    • Pure text model: Calls a visual model to generate a detailed description, then replaces the image blocks with the description text.
    • Multimodal model: Image blocks are passed to the model request as-is; the plugin does not intervene.
  3. Coverage: Replacement occurs before content is written to the session log. The plugin covers agent/pre-step (Web UI paste/drag-and-drop images, client messages) and tools/post-execute (tool results such as read_image).
  4. Capability declaration: The plugin wraps ctx.llm.resolveModelInfo / listModels, declares ['text','image'] for all providers, so capability gates such as Web upload, selectModel switching, and read_image are allowed.
  5. Error handling: If parsing one image fails, it is replaced with placeholder text only and does not interrupt the session. The original image remains saved in attachment storage.

Usage Examples

A pure text model can “understand” images through this plugin — multimodal models are passed through automatically, while pure text models trigger translation.

  1. Visual model describes the UI: The visual model translates a Windows Start menu screenshot into a detailed description.
  2. Pure text model responds: The pure text model provides a structured answer based on the description.
  3. Multiple UI translations: Screenshots of the DSH sidebar and model dropdown menu can both be accurately described.

Notes

  • Manual deployment: You need to manually deploy the files to $DSH_HOME/profiles/node_modules/deepseek-visual-plugin and add the configuration to cordis.patch.yml (ID: deepseek-visual).
  • Hot-update limitation: After modifying the plugin source code, you must restart dsh web: the web composition disables module-level HMR, and configuration hot-reloading does not reload changed plugin modules.
  • Behavior when credentials are missing: If no key is configured, the plugin still loads: messages containing images are replaced with placeholder text and produce a log warning; pure text conversations are unaffected.
  • Self-check script: verify.mjs is the functional self-check script for this plugin. Run npm test (equivalent to node verify.mjs) to verify the functionality.
  • GitHub repository: https://github.com/zhangzhimou78-code/deepseek-visual-plugin
  • Plugin directory: https://www.skillhub.cn/plugins/zhangzhimou78-code/deepseek-visual-plugin