Introduction

DeepSeek Harness (DSH) uses a plugin-based architecture. In local or on-premises deployment scenarios, text models such as DeepSeek often lack native vision capabilities, while Zhipu’s GLM models excel in visual understanding. By integrating Zhipu GLM vision models, text models such as DeepSeek can complete image-viewing tasks through the glm_vision tool.

What It Is

This plugin was developed by maintainer fightingFirefox and aims to integrate Zhipu’s GLM vision models into DSH. It mainly does three things: registers a vendor route glm-vision that connects directly to the Zhipu API; registers a tool named glm_vision, supporting local image paths and attachment IDs; and takes over the DeepSeek official route, enabling the main model to automatically invoke vision capabilities when handling images.

Installation and Activation

Install the Plugin

Ensure that the DeepSeek Harness core (@deepseek-ai/dsh) is installed and that the dsh command is available in the command line.

# 从 GitHub 直接安装
dsh plugin --profile web add github:fightingFirefox/dsh-glm-vision

After installation, the plugin is automatically added to the profile’s bundles layer.

Configure API Key

  1. Register an account at https://open.bigmodel.cn and create an API Key.
  2. Enter it on the DSH Web Settings → Models page, or directly edit $env:USERPROFILE\.dsh\.credentials.yaml:
ZAI_API_KEY: 你的智谱key(id.secret 两段式)

Restart to Take Effect

dsh web

Core Features and Usage

Routing Design

The plugin registers two main routes:
- glm-vision: Serves as the backend for the vision tool, primarily invoked by the glm_vision tool, or used as the “GLM Vision (Zhipu)” option in the model selector.
- glm_vision_router: Provides the “GLM-vision” option, supports multi-turn conversations with tool calling, and is suitable for combining pure-text GLM models with visual supplementation.

Tool Usage

Directly instruct the model in the chat to use the tool:
“Please use glm_vision to look at what is in path/to/image.png”.

Route Usage

Select a model under “GLM Vision (Zhipu)” in the model selector, such as glm-4.6v, glm-4.6v-flash, and have image-enabled conversations.

Takeover of the DeepSeek Official Route

Takeover of the deepseek-official route is enabled by default. After selecting a DeepSeek model, such as DeepSeek-V4-Pro, and uploading an image, DSH intercepts the request, converts the image into a tool-calling instruction, and makes the model call glm_vision(image='sha256:…'). GLM then returns a description, and the main model answers based on that description.

Exclusive Vision Mode

The plugin enables exclusive vision mode by default. If other vision plugins, such as vision-router, modlens, or dsh-vision, are detected, this plugin enters a dormant state and warns: “This plugin does not currently support coexistence…” If coexistence is required, disable “Coexistence Detection (Exclusive Vision)” in the settings card.

Notes

  • Dependency Environment: Requires the installed DeepSeek Harness core and a Zhipu API Key.
  • Configuration Effect: Configuration changes, such as registerTool and takeover, require restarting DSH to take effect.
  • Thinking Mode:
  • max: Deep reasoning (most accurate).
  • auto: Default mode.
  • off: Fast but less accurate; not recommended for vision tasks.
  • Version Requirement: The plugin depends on @deepseek-ai/schemastery (>=3.18.0).

References

  • Plugin Directory: https://www.skillhub.cn/plugins/fightingFirefox/dsh-glm-vision
  • GitHub Repository: https://github.com/fightingFirefox/dsh-glm-vision