Introduction¶
In development scenarios for DeepSeek Harness (DSH), if the main conversational model does not support the visual modality, sending raw image bytes can consume a large amount of context or directly cause request failures. The dsh-view-image plugin solves this problem by introducing an independent OpenAI-compatible vision model: the main conversation history retains only plain-text descriptions, image bytes do not enter the context, allowing text-only models to read images.
What is it¶
dsh-view-image is a DSH plugin maintained by johnoooooo. It allows users to configure an independent vision model to recognize images and insert the recognition results into the main conversation history as plain text, enabling text-only models to read images.
Core features¶
- Text-only models can read images: the main conversation history only retains “tool calls + plain-text results”; image bytes do not enter the model context.
- Paste and use images in Web: paste images directly into the input box. Under text-only routing, the plugin automatically persists them as persistent attachments and rewrites them as text markers.
- Inline thumbnails in chat stream: pasted images are displayed inline within user messages (right-aligned, 240px, object-fit cover). Supports click-to-zoom, a hover “Copy” button to add to the input box in one click, and dragging back to the input box.
- Terminal / headless image reading: the model can read images by passing file paths, suitable for scripted scenarios.
- UI language follows dsh: buttons and feedback text switch in real time with the dsh UI language.
- Configurable recognition behavior: default prompts, intent suffixes, and soft guidance for the main model intent can all be overridden in the profile configuration layer.
- Attachment persistence: content addressing (sha256), deduplication for identical images, integrity verification, and recoverable sessions.
- Pluggable vision model: supports OpenAI-compatible endpoints (such as Ollama, vLLM, LiteLLM, etc.).
Installation and enabling¶
Clone the repository and install it into the target profile (replace web with your profile name):
git clone https://github.com/johnoooooo/dsh-view-image.git
cd dsh-view-image
dsh plugin --profile web add file:$PWD
Restart dsh after installation to enable it.
Configuring the vision model¶
The plugin depends on an independent vision model. The configuration file is vision-model.json, located by default at ~/.dsh/vision-model.json.
Copy the example configuration file and edit it:
cp vision-model.example.opencode.json ~/.dsh/vision-model.json
Fill in the OpenAI-compatible endpoint information in the configuration file:
baseUrl: OpenAI-compatible endpoint address (such as Ollama/v1, vLLM, Alibaba Cloud Bailian, etc.).apiKey: endpoint API Key.model: vision model name.api: protocol type, defaults toopenai-completions.maxTokens: maximum output tokens.
If the configuration is missing, the plugin still loads normally but does not register the view_image tool. After completing the configuration, restart dsh or run /reload to enable it.
Usage¶
Web interface
- Paste the image directly into the input box (Ctrl+V).
- Under text-only routing, the plugin persists the image as a persistent attachment and replaces it with a text marker sent to the model—no more “current model does not support images” prompt.
- After seeing the marker, the model automatically calls
view_imageand returns a plain-text description. The chat stream displays an image thumbnail, supporting operations such as click-to-zoom, hover copy, and dragging back to the input box.
Terminal / headless
The model automatically calls the view_image tool to read the image:
dsh --profile headless "用 view_image 工具查看 /path/to/img.png 并描述内容"
Attachment storage and notes¶
Attachments are stored in ~/.dsh/attachments/v1/objects/. Files use sha256 content addressing, with deduplication for identical images.
Note: the plugin has no automatic garbage collection (GC) mechanism. Deleting a session does not clean up attachments. To clear them, manually delete that directory.