DeepSeek Harness only allows models that have declared an image input modality (inputModalities) to receive images by default. Pasting screenshots into a session with a text-only model (such as deepseek-* or glm-*) will usually fail or return an error. Existing workarounds often require switching models, using a third-party proxy, or hard-coding API keys.
The dsh-autovision plugin introduces native visual capabilities for text-only models in the DSH Web UI. At runtime, it registers a transparent Twin Provider for each text-only model, routes image requests to a configured multimodal model, and injects the transcribed text back into the context. Users can make text models “see” images without switching models or modifying settings.yaml.
Core Features¶
- Zero-friction, transparent registration and automatic routing: The plugin registers a
<provider>-autovisionTwin Provider for text-only models at runtime.agent/requestautomatically redirects requests to that Twin Provider, without manual model switching or configuration edits. - Paste-to-text: Paste an image into the editor for a text model and send it. The system automatically calls the configured multimodal model to transcribe it, then sends the generated text to the text model. The original image remains visible in the UI and is preserved in logs.
- Agent-callable tool: The plugin exposes the
autovision_read_imagetool to models. During a session, a model can proactively call this tool to read a local file (providing a path and instructions) and perform subsequent actions based on the result. - No built-in credentials: The plugin itself does not store any API keys and does not involve a relay proxy. It only calls the default visual model that you have configured in the DSH settings page (such as
minimax-m3oropencode-go). - Configuration persistence: As a pure plugin implementation, it does not modify DSH core code, and
dsh upgradewill not break the plugin’s functionality.
Install and Enable¶
Make sure your DSH Web version is at least 0.1.0-rc.6. Add the plugin by running the following command:
dsh plugin --profile web add @iroam2375/dsh-autovision
After installation, restart dsh web and go to Settings -> Plugins -> Autovision. You need to select a default visual model responsible for image-to-text transcription.
Typical Usage¶
- Paste an image directly: In any text-model session, paste an image into the editor and send it. The text model receives a faithful text transcription instead of the original image.
- Read files via the tool: Instruct the model in the conversation to read a specific file. The model can call the
autovision_read_imagetool, passing the file path and a custom instruction (such as “transcribe verbatim”), and then act based on the result.
Use Cases and Notes¶
This plugin is suitable for scenarios where you need text-only models in a DSH environment to process image information. Because the plugin runs with the permissions of the current DSH process, we recommend reviewing the source code and license before installing it.
Keep the following limitations in mind:
* First image in a new session is skipped: If the first message in a session already contains an image, that image is silently skipped. To work around this, send a plain text message before the first image message.
* File size limit: Uploaded images are limited to 5 MB.
* Transcription latency: Transcription time depends on the performance of the selected visual model (for example, minimax-m3 may take 6-9 seconds). The plugin includes an LRU cache to mitigate latency.
* UI display: In the settings page, you may see a bare provider line without an address (such as opencode-go-autovision). This is only a visual display artifact and does not affect actual functionality.
Summary¶
dsh-autovision provides a lightweight way to inject visual capabilities into text-only models in DeepSeek Harness. It does not change core configuration, does not leak keys, and supports Agent invocation. With simple configuration, you can enable text models to process image content directly.