Preface¶
Most models in DeepSeek Harness (DSH) (such as deepseek-v4-pro) are pure text. When a user pastes an image, the gateway (such as opencode-go) may return 400 unknown variant image_url. Once an image appears in the session log, every conversation replay repeats this error, causing a session deadlock. dsh-vision-guard is designed to solve this problem.
Introduction¶
dsh-vision-guard is a transparent image protection and visual analysis tool for DeepSeek Harness (dsh). It was developed by good-boy4069 and is licensed under the MIT License. The plugin uses two lines of defense to convert images into text, enabling text models to “see” images while preventing deadlocks.
Core Features¶
The plugin includes the following core capabilities:
- Convert images to text: Before an image is written into the session log, it is first converted into text by a vision model, preventing deadlocks caused by image blocks in the log.
- Repair deadlocked sessions: For sessions that are already deadlocked, the plugin rewrites images in the conversation history into OCR text during requests, restoring conversation functionality.
vision_analyzetool: Supports analysis of local OCR, PDFs,docx,pptx, and video files.- Native image passthrough: Provides native support for routes that truly support images through a whitelist mechanism.
- Zero dependencies: Built purely with Node built-in modules and has no additional complex dependencies.
Installation and Enabling¶
Install the plugin using the following command:
dsh plugin --profile web add dsh-vision-guard
After installation, the dsh web process will automatically load the plugin.
Typical Usage¶
1. Configure Vision Routing¶
In the DSH configuration, you must specify a vision model that supports images. The default configuration is as follows:
visionProvider: opencode-go
visionModel: minimax-m3
Ensure that visionModel points to a model that supports image input.
2. Use the vision_analyze Tool¶
During model invocation, you can use the vision_analyze tool. This tool selects the OCR engine through the engine parameter:
* local: Local OCR (free, character recognition).
* vision: Use the configured vision model mentioned above.
3. Check Runtime Status¶
You can check the plugin runtime status with the following code:
ctx.get('visionGuard')?.status()
4. Configure the Passthrough Whitelist¶
For routes that natively support images (such as minimax-m3), add them to the passthrough array to avoid them being rewritten into text:
passthrough:
- { provider: "opencode-go", model: "minimax-m3" }
Notes and Limitations¶
- Model dependency: When using this plugin, you must configure a model that supports images (for example,
minimax-m3). Pure-text models cannot be used as vision models. - System tool requirements: The
vision_analyzetool requires the corresponding system tools to be installed when processing documents and videos:- PDF:
pdftotext,pdfimages(poppler-utils). - Video:
ffmpeg,ffprobe. - Office documents:
python3. - Local OCR:
tesseract(optional).
- PDF:
- Image size limit: Image size is limited to within 5 MB.
- Input security: Images are treated as untrusted input and are used only as read-only data.
- Configuration coupling warning: If
input: [text, image]is declared in a text-only model configuration, this plugin must remain installed; otherwise, it may cause a deadlock.
Summary¶
dsh-vision-guard resolves compatibility issues between text models and image inputs by converting image content before it is written to the log and rewriting historical images during requests. It not only provides basic OCR capabilities but also helps prevent system deadlocks when the plugin is removed or fails.