The core idea of DeepSeek Harness (DSH) is “everything is a plugin.” When developing agents or performing model inference, the main model often supports only plain text input and cannot directly process images. The sd1g1/dsh-image-describe plugin aims to solve this pain point by injecting the describe_image tool into the model, enabling plain text models to “see” and describe images.

Plugin Introduction

This plugin is maintained by user sd1g1 and is licensed under the MIT License. It is a DeepSeek Harness host plugin whose main function is to allow a main model that does not support image input (plain text only) to support image input through the describe_image tool.

Core Features

  • Tool injection: Provides the describe_image tool, which the model can call.
  • Multiple input methods: Supports both attachmentId (references an image in the conversation) and local file paths.
  • Automatic resizing: Supports automatic image resizing (autoResize).
  • Automatic configuration: Creates the configuration file automatically on first startup.

Installation and Enablement

Install it using DSH’s plugin management command. Make sure the environment meets the dependency requirements (Node >= 20).

npx @deepseek-ai/dsh plugin --profile web add github:sd1g1/dsh-image-describe

Configuration

On first startup, the plugin automatically creates a configuration file. The path follows $DSH_HOME (usually ~/.dsh/image-describe.json). The configuration file contains settings for the vision model. If these fields are left empty, the plugin will remain dormant.

The default configuration template is as follows:

{
  "vision": { "provider": "my-vision-provider", "model": "vision-model" },
  "prompt": "请详细描述这张图片的内容",
  "maxTokens": 1500,
  "image": { "autoResize": true }
}

Configuration items:
* vision.provider / vision.model: Required. If left empty, the plugin stays dormant and does not affect normal requests.
* prompt: Optional, defaults to an image description instruction.
* maxTokens: Optional, default 1500.
* image.autoResize: Optional, default true. When enabled, preprocessing is applied to the image.

Typical Usage

After installation and configuration, the main model can call the describe_image tool.

1. Using a local image path:

describe_image({ path: '/绝对路径/图片.png' })

2. Using an attachment ID in the conversation:

describe_image({ attachment: 'attachmentId' })

Notes and Use Cases

  • Permission check: The plugin runs with the permissions of the current DSH process. It is recommended to review the source code and license before installation.
  • Format support: Supports png, jpeg, webp, and gif images.
  • Failure fallback: If processing fails, it automatically falls back to the original image, preserving basic image viewing.
  • Required fields: vision.provider and vision.model must be configured to activate the plugin.

Summary

The sd1g1/dsh-image-describe plugin provides multimodal description capabilities for DeepSeek Harness. Through simple configuration and invocation, it extends the application boundaries of plain text models.

Plugin Directory | GitHub Repository