Introduction¶
The core design philosophy of DeepSeek Harness (DSH) is “everything is a plugin.” In real-world agent development, if the main model is a text model, adding visual understanding usually requires modifying the core DSH code or performing complex system-level integration. dsh-bundle-vision aims to solve this pain point. It provides a combination of a Profile Bundle and a Plugin, and—without modifying the core DSH code—adds visual capabilities to any Profile through the Bundle patch mechanism.
What Is This¶
dsh-bundle-vision is a plugin package maintained by skillre, classified under model inference. It encapsulates the implementation of visual capabilities into an installable NPM package. Its core functionality is to register a model-facing tool named describe_image and attach this plugin to the Bundle layer of any Profile.
Core Features¶
This plugin mainly implements the following features:
- Tool registration: Registers the
describe_imagetool for LLM invocation. - Local file reading: Supports reading local image files in PNG, JPEG, WebP, and GIF formats.
- Byte submission: Submits image bytes through the attachment service.
- Routed requests: Sends LLM requests to a named multimodal routing endpoint.
- Zero core changes: All logic is implemented through plugins and Bundles, with no need to modify the core DSH code.
- Bundle Patch mechanism: Integrates via the Cordis patch mechanism.
Installation and Enablement¶
In an existing DSH environment, install it using the official command:
dsh plugin --profile <name> add dsh-bundle-vision
After installation, you need to restart the dsh <name> process. The tool will be registered to all Agents (Profile-root registration is visible across all preset scopes). The installation command automatically handles dependencies in the Bundle layer stack.
Typical Usage¶
The main model can invoke this tool directly during a conversation. When calling it, you need to provide the file path, target Provider, model name, and prompt.
Use describe_image with file_path '/path/to/photo.jpg', provider 'my-vision', model 'my-vision-model', and prompt "OCR the text in this image".
The main model specifies the Provider and Model for each invocation, enabling flexible switching without modifying Profile configuration. If a request fails, the error message clearly indicates the failure reason (for example: unsupported file extension, unsupported media type in the deployment, routing endpoint does not declare image input, file not found, or type mismatch, etc.).
Applicable Scenarios and Notes¶
Before using this plugin, pay attention to the following dependencies and configuration requirements:
- Configuration declaration: You must declare the input type for multimodal models in the
llm-pi-aisettings (for example, includeimagein theinputfield). - Version dependencies: The required peerDependencies must satisfy
>=0.1.0-rc.6. - Conversation content: Tool results are plain text and do not introduce image blocks into the conversation, so a plain-text main model can use them directly.
- Runtime dependencies: The plugin depends on
ctx.fs(file system capabilities),ctx.attachments(attachment service), andctx.llm(LLM invocation).
Short Ending¶
dsh-bundle-vision brings visual capabilities to DeepSeek Harness without core code changes, by combining tool registration and the Bundle patch mechanism. It leverages the existing ctx context and attachment service to allow text models to call external multimodal routing endpoints for image processing. For more details, refer to the Catalog page or the GitHub repository.