In a DSH environment, directly pasting images often causes conversation errors or prevents images from rendering properly in the chat window, which can interrupt the flow of dialogue. The auto-vision plugin solves this problem by automatically cleaning image data and registering the see_image tool. It allows the text model to automatically invoke a vision model to read content when needed, while preserving the original dialogue logic.

Plugin Positioning

This is a DSH plugin designed to provide integrated image handling. It addresses the issues of “pasted images causing errors” and “abnormal image display.” It supports free vision models from Zhipu and the ModelScope community, and is also compatible with any OpenAI-compatible platform.

Installation and Enablement

  1. Copy this directory to the node_modules directory under the DSH profile (for example, ~/.dsh/profiles/web/node_modules/auto-vision/).
  2. Insert the following content into the profile’s cordis.patch.yml:
    - insert:
        - id: auto-vision
          name: auto-vision
  1. After configuring the Token, restart the DSH process.

Configure the Token

The plugin supports configuration through environment variables or a credentials file. Supported vision platforms include:

  • Zhipu BigModel: Set ZHIPU_API_KEY; the model is GLM-4V-Flash.
  • ModelScope: Set MODELSCOPE_API_KEY; the model is Qwen3-VL.
  • Other OpenAI-compatible platforms: Set VISION_API_KEY, VISION_ENDPOINT, and VISION_MODEL.

You can also write them directly to the ~/.dsh/.credentials.yaml file.

export ZHIPU_API_KEY=xxx          # 用智谱
export MODELSCOPE_API_KEY=xxx     # 用魔搭

Switch Vision Source

ModelScope is used by default. Switch to Zhipu by setting the environment variable VISION_PROVIDER=zhipu or by setting provider: zhipu in the plugin config.

For local or other platforms, set endpoint and models in the plugin config:

config:
  endpoint: http://127.0.0.1:11434/v1/chat/completions
  models: [qwen2.5-vl:7b]

Notes

  • Restart the DSH process for changes to take effect.
  • Requires @deepseek-ai/dsh-tools >= 0.0.1-rc.0.
  • Requires Node.js version >= 18.
  • License: MIT.

Conclusion

The plugin makes DSH more stable when handling multimodal conversations, without requiring manual intervention in the image cleaning process. For more details, refer to the plugin directory or the GitHub repository.