Preface¶
The core philosophy of DeepSeek Harness (DSH) is “everything is a plugin,” aiming to extend capabilities through modular expansion. A pure text large language model (the “brain”) cannot process pixel data directly, making it prone to guessing or failing in visual scenarios. jolly-dsh-vision is a ModLens-style vision bridging plugin that solves this problem by introducing a dedicated vision model (the “eyes”). It enables pure text models to understand images and provides structured evidence output.
Core Capabilities¶
This plugin mainly provides two visual pathways:
visiontool: The brain can actively invoke this tool, passing a local path or URL. The plugin stores the image in the attachment store, calls the vision model through the LLM Seam, and returns structured evidence JSON containing fields such assummary,ocr, andlayout. The brain answers based on this evidence rather than guessing.(ds vision)vision twin model: Entries with the(ds vision)suffix appear in the model selector (for example,DeepSeek-V4-Pro (ds vision)). After selecting such an entry, you can paste or drag an image directly into the input box. Before the request is sent, the plugin automatically converts the image into evidence text and injects it into the brain.
Installation and Enablement¶
The installation process requires the DSH command-line tool.
dsh plugin --profile web add jolly-dsh-vision
After installation, ensure dsh.profile.bundles includes jolly-dsh-vision, and restart dsh web to load the plugin.
Prerequisites¶
- Environment requirement: Node.js version must be 20 or higher.
- Configure the brain and the eyes: Configure two models in
~/.dsh/settings.yaml. One acts as the brain (pure text), and the other acts as the eyes (vision).
agent-default-model:
provider: deepseek-official
model: deepseek-v4-pro # Brain
llm-deepseek:
models:
- id: deepseek-v4-flash-vision-exp
name: DeepSeek-V4-Flash-Vision-Exp
inputModalities:
- text
- image # The eyes must declare image input support
Typical Usage¶
Pathway A: Using the vision tool¶
During a conversation, if the brain needs to look at an image, it will automatically invoke the vision tool. You can pass a local file path or an HTTP(S) URL directly.
Pathway B: Direct image pasting and the (ds vision) twin¶
- Find the entry with the
(ds vision)suffix in the model selector. - Paste or drag an image directly into the input box.
- The image is automatically converted into evidence text, and the brain references this evidence when answering. The same image in the same session is converted only once.
Note: Regular pure text model entries do not accept pasted images. You must select an entry with the
(ds vision)suffix.
Security and Privacy¶
- Zero API Key contact: The plugin does not read, record, or forward any API Key. Credentials are resolved by the DSH Harness credential seam for each call, and the code does not construct Authorization headers.
- SSRF protection: The
visiontool rejects private network ranges, loopback addresses (such aslocalhost), and link-local addresses by default, preventing malicious prompts from inducing requests to internal resources. If internal resources are required, you can enableallowPrivateUrlsin the configuration. - Trust boundary: The plugin runs with the current DSH process privileges, and the
visiontool can read any image file readable by the process. Image paths (including resolved absolute paths) will appear in session logs. Be mindful of data leakage risks when sharing sessions among multiple people.
Known Limitations¶
- Format support: Only PNG, JPEG, GIF, and WebP are supported. HEIC files must be transcoded first.
- Routing mechanism: The vision twin is currently a single-routing implementation and does not yet support automatic multi-routing discovery like modlens.
- Dependencies: Zero dependencies (pure Node.js built-in modules).
Conclusion¶
jolly-dsh-vision adds vision capabilities to DSH with a highly minimal architecture. By using deepseek-v4-flash-vision-exp as the eyes, it allows brain models such as deepseek-v4-pro to perform more accurate visual reasoning based on structured evidence. The project code is licensed under the MIT license and derived from liustack/modlens.