Introduction¶
The philosophy of DeepSeek Harness is “everything is a plugin.” When building agents, pure text models cannot directly handle image inputs. This plugin solves this problem: it intercepts messages containing images and determines whether the current model natively supports image. Multimodal models are passed through directly, while pure text models call a visual model to generate a description and then replace image blocks with text. This way, pure text models can also “understand” images.
Installation and Enablement¶
This plugin is deployed to the Harness home directory as a single-file ESM plugin and enabled through a home-level patch. It applies to all profiles (web / headless, and any newly created profiles in the future).
Windows (PowerShell)
# 1. 部署插件代码
$dest = "$env:USERPROFILE\.dsh\profiles\node_modules\deepseek-visual-plugin"
New-Item -ItemType Directory -Path $dest -Force | Out-Null
Copy-Item index.js, package.json, cordis.patch.yml -Destination $dest -Force
# 2. 在 $DSH_HOME/cordis.patch.yml 中加入(不存在则创建)
# - insert:
# - id: deepseek-visual
# name: deepseek-visual-plugin
# config:
# model: <视觉模型名>
# baseUrl: <OpenAI 兼容端点>
# apiKeyEnv: <凭证引用名>
Linux / macOS (Bash / Zsh)
# 1. 部署插件代码
dest="$HOME/.dsh/profiles/node_modules/deepseek-visual-plugin"
mkdir -p "$dest"
cp index.js package.json cordis.patch.yml "$dest"
# 2. 在 $DSH_HOME/cordis.patch.yml 中加入
# - insert:
# - id: deepseek-visual
# name: deepseek-visual-plugin
# config:
# model: <视觉模型名>
# baseUrl: <OpenAI 兼容端点>
# apiKeyEnv: <凭证引用名>
It can also be installed as a regular bundle into a single profile:
dsh plugin --profile web add C:\path\to\deepseek-visual-plugin
Configuration¶
The plugin is configured through cordis.patch.yml. The adjustable parameters are as follows:
| Field | Default | Description |
|---|---|---|
model |
qwen3.7-plus |
Visual model name (model ID on an OpenAI-compatible endpoint) |
baseUrl |
DashScope-compatible endpoint | When left empty, falls back to QWEN_BASE_URL / DASHSCOPE_BASE_URL / OPENAI_BASE_URL |
apiKey |
Empty | Explicit API key override (optional; writing it in plaintext configuration is not recommended) |
apiKeyEnv |
Empty | Credential reference name (e.g., DOUBAO_API_KEY), resolved through dsh’s credentials service |
imagePrompt |
Built-in Chinese prompt | Image description prompt |
timeoutMs |
60000 |
Timeout for a single visual request |
API Key resolution priority: config.apiKey (explicit override) → config.apiKeyEnv (credentials service resolving ~/.dsh/.credentials.yaml) → environment variable fallback (QWEN_API_KEY / DASHSCOPE_API_KEY / OPENAI_API_KEY).
Recommended practice: store the key in ~/.dsh/.credentials.yaml and write only the apiKeyEnv reference in the configuration, avoiding plaintext keys in the configuration file.
# ~/.dsh/.credentials.yaml
DOUBAO_API_KEY: <你的 key>
# cordis.patch.yml
- insert:
- id: deepseek-visual
name: deepseek-visual-plugin
config:
model: doubao-seed-2-0-mini-260428
baseUrl: https://ark.cn-beijing.volces.com/api/v3
apiKeyEnv: DOUBAO_API_KEY
How It Works¶
The plugin is implemented against the native protocol and includes the following main logic:
- Interception and detection: Intercepts messages containing images and checks whether the current model natively supports image.
- Routing:
- Pure text model: Calls a visual model to generate a detailed description, then replaces the image blocks with the description text.
- Multimodal model: Image blocks are passed to the model request as-is; the plugin does not intervene.
- Coverage: Replacement occurs before content is written to the session log. The plugin covers
agent/pre-step(Web UI paste/drag-and-drop images, client messages) andtools/post-execute(tool results such asread_image). - Capability declaration: The plugin wraps
ctx.llm.resolveModelInfo/listModels, declares['text','image']for all providers, so capability gates such as Web upload,selectModelswitching, andread_imageare allowed. - Error handling: If parsing one image fails, it is replaced with placeholder text only and does not interrupt the session. The original image remains saved in attachment storage.
Usage Examples¶
A pure text model can “understand” images through this plugin — multimodal models are passed through automatically, while pure text models trigger translation.
- Visual model describes the UI: The visual model translates a Windows Start menu screenshot into a detailed description.
- Pure text model responds: The pure text model provides a structured answer based on the description.
- Multiple UI translations: Screenshots of the DSH sidebar and model dropdown menu can both be accurately described.
Notes¶
- Manual deployment: You need to manually deploy the files to
$DSH_HOME/profiles/node_modules/deepseek-visual-pluginand add the configuration tocordis.patch.yml(ID: deepseek-visual). - Hot-update limitation: After modifying the plugin source code, you must restart
dsh web: the web composition disables module-level HMR, and configuration hot-reloading does not reload changed plugin modules. - Behavior when credentials are missing: If no key is configured, the plugin still loads: messages containing images are replaced with placeholder text and produce a log warning; pure text conversations are unaffected.
- Self-check script:
verify.mjsis the functional self-check script for this plugin. Runnpm test(equivalent tonode verify.mjs) to verify the functionality.
Links¶
- GitHub repository: https://github.com/zhangzhimou78-code/deepseek-visual-plugin
- Plugin directory: https://www.skillhub.cn/plugins/zhangzhimou78-code/deepseek-visual-plugin