DeepSeek Harness (DSH) adopts a plug-in architecture designed to extend the system’s native capabilities. Pure text large models cannot directly “see” images and require external tool assistance. dsh-siliconflow-vision is a plug-in that registers an analyze_image tool, sends image input to SiliconFlow’s vision model (using Qwen/Qwen3-VL-32B-Instruct by default), and returns the recognition results to the DSH session, enabling the conversational model with complete image analysis capabilities.
This is a model inference plug-in maintained by ShiXiangYu2 and licensed under the MIT License.
Core Features¶
This plug-in primarily provides the following visual recognition capabilities:
- Recognize server-local image files (passed via file paths).
- Recognize HTTP(S) image URLs.
- Recognize Base64 Data URLs.
- Custom recognition instructions (Prompts), such as “Recognize all text in the image” or “What animals are in the image”.
- Support switching to different vision models (Qwen3-VL series, GLM-4.5V, PaddleOCR-VL, etc.).
- Provide an optional interactive “paste and recognize” panel (dynamic plug-in form).
Installation and Activation¶
Use the dsh plugin command to install the plug-in into the specified profile.
dsh plugin --profile demo add ./dsh-siliconflow-vision
After installation, load this profile when starting DSH, and the analyze_image tool will automatically become active.
Configuring the API Key¶
The plug-in code does not hardcode API keys. You need to configure SiliconFlow access credentials via an environment variable or a key file.
Option 1: Environment Variable
export SILICONFLOW_API_KEY="sk-xxxxxxxx"
Option 2: Key File
Create a plain text file in any of the following locations, with the API key as its content:
$DSH_HOME/siliconflow.key
~/.dsh/siliconflow.key
The API key must be obtained from the SiliconFlow console.
Typical Usage¶
In a DSH conversation, simply enter an instruction that includes an image path or URL. The model will automatically invoke the analyze_image tool.
Example:
分析一下 /root/data/photo.png 里有什么
识别这张图的文字:https://example.com/screenshot.png
The tool call parameters are as follows:
image: image source (required), supporting local paths, HTTP URLs, or Data URLs.prompt: instruction for the image; if omitted, a generic Chinese description is used.model: specified SiliconFlow model ID, defaulting toQwen/Qwen3-VL-32B-Instruct.maxTokens: output token limit, defaulting to 1024.
Optional Model List¶
Depending on the recognition requirement, you can switch to different models:
| Model ID | Features |
|---|---|
Qwen/Qwen3-VL-32B-Instruct |
Default, strong recognition capability |
Qwen/Qwen3-VL-8B-Instruct |
Faster, more resource-efficient |
Qwen/Qwen3-VL-30B-A3B-Instruct |
Cost-effective choice |
zai-org/GLM-4.5V |
General-purpose visual understanding |
PaddlePaddle/PaddleOCR-VL-1.5 |
Focused on OCR text recognition |
Technical Notes and Precautions¶
- Runtime Environment: Node.js >= 18 is required.
- Request Format: Uses an OpenAI-compatible API by sending a POST request to
https://api.siliconflow.cn/v1/chat/completions. - Local Image Processing: After reading a local image, it is directly converted into a base64 data URL and passed in. No temporary disk file is created, avoiding permission and path issues.
- Security: The plug-in runs with the permissions of the current DSH process. Before using it, it is recommended to review the source code to confirm that its behavior is as expected.
This plug-in connects SiliconFlow’s vision capabilities, giving DSH sessions the ability to “see” images. For more technical details and source code, visit the GitHub repository.