Introduction¶
DeepSeek Harness (DSH) extends functionality through a plugin mechanism. In multimodal conversation scenarios, users often need to upload images and manually describe their content, which increases operational cost and may lead to imprecise descriptions. dsh-vision-plugin solves this problem by automatically transcribing image content and injecting descriptions before a message is sent, lowering the barrier to multimodal interaction.
Plugin Overview¶
This plugin is maintained by developer zcma11 and is categorized under model inference. Its core logic is as follows: when a user uploads or pastes an image in the chat box, upon sending, the plugin transcribes the image content using a vision model (DashScope) or offline Windows OCR, and automatically inserts the description at the beginning of the message.
Installation and Enablement¶
Install it using the official DSH plugin management command:
npx @deepseek-ai/dsh plugin --profile web add dsh-vision-plugin
Note: If you encounter the ERR_PNPM_ADDING_TO_ROOT error while using pnpm 9, append the --workspace-root flag to the end of the command.
Usage¶
Uploading and Pasting Images¶
In the chat input box toolbar, click the corresponding icon to select an image, or paste directly with Ctrl+V. Supported formats include screenshots (such as QQ screenshots and system screenshots). BMP images are automatically converted to PNG.
Configuring the Vision Model¶
Since pure OCR mode does not depend on an external model, only vision transcription mode requires configuration. Go to Settings -> General -> Vision Transcription Model, select an image-input capable model from the dropdown (for example, dashscope/qwen3.6-plus), and ensure that DASHSCOPE_API_KEY is configured.
Sending and Transcription¶
-
Select a mode:
- Click the 📝 (Pure OCR) button: calls only the built-in Windows OCR engine and runs offline.
- Click the 🖼️ (Vision Transcription) button or paste an image: calls only the vision model for transcription. -
Sending process:
- During transcription, the preview bar above the input box displays “⏳ Parsing image…”, and new send requests are blocked.
- After successful transcription, the message content automatically begins with【解析了提供图片,图片内容是<描述>】, followed by your original text. The preview bar then clears automatically.
- If transcription fails, the message is not sent. A red notification appears at the top, clearly distinguishing whether it was a “Vision model invocation failed” or “OCR recognition failed” error.
Applicable Scenarios and Cautions¶
- Applicable scenarios: Suitable for developers or users who need multimodal conversation in DSH.
- System dependencies: Pure OCR functionality requires Windows 10+ and PowerShell 5.1, and the plugin has been verified in a Windows environment.
- Security and storage: Images are stored only in memory as base64 and are not written to disk, meeting security isolation requirements.
- Mode routing: The plugin uses strict routing. The OCR button uses only OCR, while the vision button/paste uses only the vision model, with no automatic fallback.
Conclusion¶
dsh-vision-plugin simplifies the multimodal interaction workflow through automated image parsing and description injection. See the GitHub repository for more details and source code.