Model inference in DSH typically relies on adapters to connect with various large models. For pure-text LLMs (such as DeepSeek R1, GPT-4 base, etc.), directly pasting or dragging an image cannot be recognized by the model; it is usually ignored or causes the request to fail. The dsh-plugin-image-input plugin takes over the image input flow. By calling a vision API locally, it converts the image into a structured text description, enabling pure-text models to obtain image context.

Plugin Information

  • Name: dsh-plugin-image-input
  • Maintainer: Elohia
  • Category: Model Inference
  • License: MIT
  • Repository: https://github.com/Elohia/dsh-plugin-image-input
  • Directory Page: https://www.skillhub.cn/plugins/Elohia/dsh-plugin-image-input

Installation and Enablement

Run the following command in a terminal to complete the installation. If it is already installed, no repeat is needed.

# 方式 A:从本地目录安装(拿到插件目录后)
dsh plugin --profile web add D:\path\to\dsh-plugin-image-input

# 方式 B:从 npm / GitHub 安装
dsh plugin --profile web add dsh-plugin-image-input

After installation is complete, restart DSH web for it to take effect.

Configuring the Vision API

The plugin relies on an external vision API to convert images to text and needs to be configured once in the settings page. The configuration file is stored at ~/.config/mm-vision/config.json, and environment variables can also be used as a fallback.

  1. Open the DSH web settings page.
  2. Find the Image to Text configuration item.
  3. Fill in the following fields:
Field Description Example
baseUrl OpenAI-compatible endpoint address (does not include /chat/completions) https://dashscope.aliyuncs.com/compatible-mode/v1
model Vision model name qwen-vl-max / gpt-4o / glm-4v
apiKey Your API Key (leave empty to keep unchanged) sk-...
maxTokens Maximum output tokens 2048

Supported environment variables include MM_VISION_API_KEY, DASHSCOPE_API_KEY, QWEN_API_KEY, OPENAI_API_KEY, and GEMINI_API_KEY.

Core Features

The plugin provides the following capabilities:

  1. Automatic Text Conversion: After pasting or dragging an image into the conversation input box, simply press Enter or click Send. The plugin automatically converts the image into a structured text description and sends it.
  2. Structured Description: The generated text includes information such as canvas, elements, and percentage coordinates, making it suitable for scenarios like K-line charts, screenshots, and other charts.
  3. Multimodal Compatibility: If the model being used supports a native image channel (such as qwen-vl or gpt-4o), the plugin automatically allows it to pass through without interception.
  4. Manual Preview: Clicking the 🖼️ button on the left side of the input box converts the image to text and inserts it into the input box so the user can edit it before sending.
  5. General Compatibility: It does not depend on DSH model adapters and supports any OpenAI-compatible vision endpoint.

Usage

  1. Input an Image: Paste or drag an image into the conversation input box (a thumbnail preview will be shown).
  2. Send for Processing:
    • Press Enter or click the Send button directly. The plugin converts the image into a text description and sends it together with your text.
    • Or click the 🖼️ button first to insert only the description into the input box, edit it, and then send.
  3. Wait for Generation: During conversion, the input box displays a status. Wait about 1 minute (complex charts may be slower).
  4. View the Result: The model receives “your text + image recognition context” and no longer rejects the message because of the image.

Technical Notes and Cautions

  • Process Permissions: The plugin runs a node subprocess with fixed content under the danger-full-access policy. Check the source code and license before installation.
  • Security Restrictions: Local page calls only (Origin validation); requests are sent only to the vision API address you configure.
  • Error Handling: If conversion fails, the plugin shows an error but retains the image, so input content is not lost.
  • Multiple Image Support: It supports processing multiple images at the same time; the plugin converts them one by one and sends them together.

Summary

dsh-plugin-image-input addresses the pain point that pure-text LLMs cannot directly process image input. With simple pasting and sending, the model can receive a contextual description of the image, thereby improving the accuracy of the conversation.