Official models of DeepSeek Harness (DSH), such as deepseek-v4-flash and deepseek-v4-pro, are text-only. Their API endpoints do not accept image bytes, and their model metadata also declares support for text input only. In the DSH environment, this means screenshots cannot be pasted into sessions that use these models, and the built-in read_image tool cannot be used either.
dsh-vision-bridge aims to solve this problem. It uses a locally running vision model to allow text-only models to receive and understand images.
What This Is¶
This is a DeepSeek Harness plugin maintained by YuLee-314. It uses the qwen2.5vl model running in local Ollama to handle vision tasks, without requiring cloud API keys. All image processing is completed locally, and no image bytes are sent to the DeepSeek API.
Core Features¶
The plugin implements its functionality across three levels:
- Native image experience: Text-only models can also handle images like vision models. Pasting a screenshot generates a thumbnail and image blocks, and the
read_imagetool works normally. - Request-layer transparency: There is no prompt hacking and no predefined branching. The plugin intercepts image blocks at the adapter layer, converts them into text descriptions, and then sends them to the DeepSeek API.
- Structured output and coordinates: The vision tools support generating element bounding boxes with coordinates (
[0,1000]normalized coordinates), and support region cropping, dual-image comparison, and clipboard reading. - Nine vision tools: Provides
describe_image(description),extract_text(OCR),structured_scan(structured scanning),query_region(region query),detect_elements(element detection),locate_object(object localization),compare_images(dual-image comparison),read_clipboard(clipboard reading), andcheck_health(health check). - Duplicate image caching: Duplicate images hit a content-hash cache and do not require repeated inference.
- Self-contained and coexisting: Single-file distribution (approximately 29 KB), with no machine-specific path dependencies. The plugin coexists with the official route, with the official route as a fallback; the plugin decides, based on model metadata, whether to preserve the native image stream or convert to a local path.
Installation and Enablement¶
The installation command is:
dsh plugin --profile web add
After installation, the plugin registers a “twin” route named deepseek-vision and injects the nine vision tools into DSH’s toolset.
Typical Usage¶
- Paste images: Paste a screenshot in an image-supporting session, and DSH automatically generates a thumbnail and image blocks. If using the native text route, the image is saved as a local path for the vision tools to read.
- Local visual processing: Use the local
qwen2.5vlmodel to generate image descriptions for DeepSeek to use. - Interactive tools: Interact using tools such as
read_clipboardorcompare_images. The model can call these tools to analyze images. - Twin Route: Process native image streams via the
deepseek-visionroute.
Requirements and Dependencies¶
- Node.js: Version must be 22.19 or higher.
- Local model: Ollama must be running locally, with the
qwen2.5vlmodel loaded. - Dependencies: Uses
sharpfor image processing andopenaias the API wrapper.
Notes¶
- The plugin runs with the permissions of the current DSH process. Review the source code and license (MIT) before installation.
- All image processing is completed locally, ensuring privacy and security.