DeepSeek Harness (dsh) is a plugin-centric agent development framework. When building agents, if a pure text model that does not support vision is used, uploading images directly can interrupt processing. The dsh-mingmu plugin adds visual processing capability to such models: when it detects that the model cannot process images, it automatically calls a vision model to recognize the content and feeds the recognized text back to the main model.
This plugin is maintained by Lab-sku. It is a pure ESM plugin, does not modify source code, and can be restored after uninstallation.
Core Features¶
- Triggered by model capability: At runtime, the plugin intercepts model calls and only bridges models that do not support visual input; models natively supporting vision keep the native processing path.
- Vision model recognition: Built-in heuristic recognition rules for global vision models, supporting 20+ model families including OpenAI, Anthropic, Google, Qwen, MiniMax, Kimi, GLM, DeepSeek-VL. If automatic recognition fails, it will still attempt bridging.
- Automatic upgrade strategy: Defaults to a cascade strategy; when the primary recognition model fails or returns an empty result, it automatically switches to the backup model. Advanced users can enable concurrent race mode via the environment variable
DSH_VISION_STRATEGY=race. - Visualized configuration: On the dsh Web “Settings → Plugins” page, provider, model, and strategy can be configured directly, with changes taking effect in real time.
- No silent failures: Provides two modes:
annotate(feeds the failure reason as text to the main model) anderror(strictly reports an error), avoiding silently dropping images. - Credential security: API keys are stored in the official dsh credential store (
~/.dsh/.credentials.yaml) and are not written to plugin configuration files. - Environment variable fallback: In addition to the Web settings page, all key parameters can also be configured via environment variables.
Installation¶
Install the plugin using the official command:
dsh plugin --profile web add github:Lab-sku/dsh-mingmu
Configuration and Usage¶
Option 1: Web Settings Page (Recommended)¶
Configure it on the dsh Web “Settings → Plugins → Mingmu” page. Changes take effect immediately after saving.
Option 2: Environment Variables¶
Configure the following key items in environment variables:
| Variable | Default | Description |
|---|---|---|
DSH_VISION_ENABLED |
true |
Main switch |
DSH_VISION_MODEL |
Qwen/Qwen3-VL-8B-Instruct |
Primary visual recognition model |
DSH_VISION_MODEL_UPGRADE |
Qwen/Qwen3-VL-32B-Instruct |
Backup recognition model |
DSH_VISION_STRATEGY |
cascade |
Strategy: cascade or race |
DSH_VISION_BASE_URL |
https://api.siliconflow.cn/v1 |
Vision API endpoint |
DSH_VISION_API_KEY_ENV |
SILICONFLOW_API_KEY |
API key environment variable reference |
DSH_VISION_FAILURE_MODE |
annotate |
Failure mode: annotate or error |
DSH_VISION_MAX_TOKENS |
4096 |
Maximum tokens for recognition |
DSH_VISION_TIMEOUT_MS |
180000 |
Recognition timeout in milliseconds |
API Key Management¶
The plugin supports two ways to configure API keys. Use either one:
- Web UI: Paste the key into the “Vision API Key” input field on the Mingmu settings page, then click Save. The key is written to the official dsh credential store.
- Manual entry: Manually add it to the
~/.dsh/.credentials.yamlfile. For example:
SILICONFLOW_API_KEY: sk-xxx
Notes¶
- Dependency version: This plugin is tested with dsh
0.1.0-rc.6. The dsh interface is still in the rc stage, and future version upgrades may require adaptation. - Performance overhead: Because model capability must be checked, each step for image-containing messages makes one additional
resolveModelInfocall. Be aware of latency-sensitive scenarios. - Privacy and endpoint: By default, images are sent to SiliconFlow (
https://api.siliconflow.cn/v1). Users can replace it with any OpenAI-compatible endpoint, but must evaluate privacy and availability themselves. - Node.js version: The runtime requires Node.js >= 22.19.0.
- Concurrency strategy:
race/cascadeonly take effect within a single session and do not change dsh’s own concurrency model logic.
Summary¶
dsh-mingmu intercepts and bridges model calls to solve the issue where pure text models cannot process images in the dsh environment. It provides a visual configuration entry point and environment variable support, and places emphasis on API key security. Before use, it is recommended to confirm the local environment versions (Node.js and dsh) and evaluate the privacy policy of third-party API endpoints.