Introduction¶
In the DeepSeek Harness (DSH) ecosystem, text-only primary models (such as DeepSeek model series) cannot directly process images. When handling UI review, OCR, or image-based QA, you usually need to start a separate vision subagent, which introduces additional context overhead and state-management complexity.
dsh-vision-resident solves this problem. It is a vision assistant plugin that resides in the host process. Instead of relying on an extra subagent process, it mounts hooks and components so that text-only primary models can also “see” images.
What It Is¶
This is a memory-type plugin for DeepSeek Harness, maintained by AnnanRen under the MIT license.
Its core value lies in:
1. Zero subagent overhead: The vision model is connected directly (via ctx.llm.stream). Each recognition is a single LLM inference, with no agent context initialization cost.
2. Persistent conversation: The vision subagent retains continuous context, remembers the images it has seen and the recognition results, and does not require resending images during follow-up questions.
3. Session isolation: Each DSH conversation has independent vision history, context, and counters, preventing cross-contamination between sessions.
4. Persistence and cleanup: Recognition history and transcription caches are persisted to disk and survive restarts. Archival linkage and idle cleanup are supported.
Core Features¶
Tools and Hooks¶
The plugin provides two tools, describe_image and subagent_vision, for reading images from disk and converting them into text results. It also uses the llm/stream hook to implement automatic image transcription on paste: when the primary model does not support images, the plugin automatically calls the vision model to convert the image into text and injects it into the request.
Status Capsule and Details Page¶
To the left of the input box, a status capsule displays the working state, model name, and recognition count of the vision subagent in real time. Clicking the capsule opens a collapsible details page where you can view the recognition history for each image (prompt, thinking process, and output). It can be dismissed with Esc or by clicking the background.
Session Management¶
Vision history is isolated per DSH conversation and retains the most recent 10 recognition rounds by default. The session limit is 40 entries (LRU policy), and records idle for more than 30 days are automatically cleaned up.
Installation and Enablement¶
Installation requires manually cloning the repository and mounting it into a DSH profile.
- Clone the repository locally:
git clone https://github.com/AnnanRen/dsh-vision-resident.git
- Create a symbolic link and mount it into
node_modules:
ln -s "$PWD/dsh-vision-resident" ~/.dsh/profiles/node_modules/dsh-vision-resident
- Edit
~/.dsh/profiles/web/cordis.patch.ymland add the following to theinsertlist:
- id: dsh-vision-resident
name: dsh-vision-resident
- Restart the DSH Web service.
Configuration¶
Vision Model Settings¶
By default, it uses opencode-go/mimo-v2.5 as the primary model and opencode-go/minimax-m3 as the fallback. You can override these via environment variables:
export VISION_PROVIDER=opencode-go
export VISION_MODEL=mimo-v2.5
export VISION_FALLBACK_MODELS=opencode-go/minimax-m3
Image Allowance for Text-Only Primary Models¶
The Harness checks the input declaration of the selected model. To allow a text-only model (such as deepseek-v4-flash) to support pasted images, declare the following in ~/.dsh/settings.yaml:
input: [text, image]
For models in the built-in force-convert list (such as deepseek-v4-flash / deepseek-v4-pro), the plugin forcibly transcribes the image instead of sending the original image directly to the upstream.
Usage¶
- Read an image from disk:
Use thedescribe_imageorsubagent_visiontool with theimage_pathparameter. For example:
await ctx.callTool('describe_image', { image_path: '/path/to/screenshot.png' });
-
Text-only follow-up:
Omit theimage_pathparameter and ask a follow-up question in plain text based on the context already known to the vision subagent. -
Custom models:
Switch vision models or configure the fallback chain using the environment variables above.
Known Limitations¶
- Model switching limitation: If the conversation already contains images, switching to a model that does not support images is rejected by the Harness (
model-unavailable), and you must start a new conversation to switch. - Tool stripping: When the current prompt contains images, tools are stripped (to prevent weaker models from reflexively searching for files). Perform operations step by step.
- Historical context: Historical context includes text only and does not include historical images. Multi-image requests may be unstable on some endpoints, and historical recognition results are already complete text descriptions.
- Session cleanup: After the main conversation is archived, the corresponding vision records are cleared within 30 seconds.
Conclusion¶
By integrating vision capabilities into the main process flow, dsh-vision-resident provides a lightweight vision enhancement solution for DSH. It is well suited for scenarios that require frequent vision tasks but do not want to introduce a complex multi-agent architecture.
For more details and source code, visit the GitHub repository.