DeepSeek Harness (DSH) adopts a plug-in architecture designed to extend the system’s native capabilities. Pure text large models cannot directly “see” images and require external tool assistance. dsh-siliconflow-vision is a plug-in that registers an analyze_image tool, sends image input to SiliconFlow’s vision model (using Qwen/Qwen3-VL-32B-Instruct by default), and returns the recognition results to the DSH session, enabling the conversational model with complete image analysis capabilities.

This is a model inference plug-in maintained by ShiXiangYu2 and licensed under the MIT License.

Core Features

This plug-in primarily provides the following visual recognition capabilities:

  • Recognize server-local image files (passed via file paths).
  • Recognize HTTP(S) image URLs.
  • Recognize Base64 Data URLs.
  • Custom recognition instructions (Prompts), such as “Recognize all text in the image” or “What animals are in the image”.
  • Support switching to different vision models (Qwen3-VL series, GLM-4.5V, PaddleOCR-VL, etc.).
  • Provide an optional interactive “paste and recognize” panel (dynamic plug-in form).

Installation and Activation

Use the dsh plugin command to install the plug-in into the specified profile.

dsh plugin --profile demo add ./dsh-siliconflow-vision

After installation, load this profile when starting DSH, and the analyze_image tool will automatically become active.

Configuring the API Key

The plug-in code does not hardcode API keys. You need to configure SiliconFlow access credentials via an environment variable or a key file.

Option 1: Environment Variable

export SILICONFLOW_API_KEY="sk-xxxxxxxx"

Option 2: Key File

Create a plain text file in any of the following locations, with the API key as its content:

$DSH_HOME/siliconflow.key
~/.dsh/siliconflow.key

The API key must be obtained from the SiliconFlow console.

Typical Usage

In a DSH conversation, simply enter an instruction that includes an image path or URL. The model will automatically invoke the analyze_image tool.

Example:

分析一下 /root/data/photo.png 里有什么
识别这张图的文字:https://example.com/screenshot.png

The tool call parameters are as follows:

  • image: image source (required), supporting local paths, HTTP URLs, or Data URLs.
  • prompt: instruction for the image; if omitted, a generic Chinese description is used.
  • model: specified SiliconFlow model ID, defaulting to Qwen/Qwen3-VL-32B-Instruct.
  • maxTokens: output token limit, defaulting to 1024.

Optional Model List

Depending on the recognition requirement, you can switch to different models:

Model ID Features
Qwen/Qwen3-VL-32B-Instruct Default, strong recognition capability
Qwen/Qwen3-VL-8B-Instruct Faster, more resource-efficient
Qwen/Qwen3-VL-30B-A3B-Instruct Cost-effective choice
zai-org/GLM-4.5V General-purpose visual understanding
PaddlePaddle/PaddleOCR-VL-1.5 Focused on OCR text recognition

Technical Notes and Precautions

  • Runtime Environment: Node.js >= 18 is required.
  • Request Format: Uses an OpenAI-compatible API by sending a POST request to https://api.siliconflow.cn/v1/chat/completions.
  • Local Image Processing: After reading a local image, it is directly converted into a base64 data URL and passed in. No temporary disk file is created, avoiding permission and path issues.
  • Security: The plug-in runs with the permissions of the current DSH process. Before using it, it is recommended to review the source code to confirm that its behavior is as expected.

This plug-in connects SiliconFlow’s vision capabilities, giving DSH sessions the ability to “see” images. For more technical details and source code, visit the GitHub repository.