Introduction

The core design philosophy of DeepSeek Harness (DSH) is “everything is a plugin”. When developing text-model-based agents, the model itself often lacks the direct capability to process images or audio, or generate images. These capabilities need to be achieved by calling external tools. The wuwangmao/dsh-qwen-multimodal plugin provides a unified entry point that bridges the multimodal capabilities of the Qwen series for use by pure text models in DSH.

What Is This

This is a DSH bundle designed to provide multimodal support for DeepSeek Harness. It uses Qwen’s API (including vision-language models, speech recognition, and image generation models) to enable text models to understand images, transcribe audio, and generate images locally.

The plugin is a pure JavaScript (Cordis) bundle. At runtime, it depends on the host-mounted subprocess or tools service and does not require build steps.

Core Capabilities

The plugin provides three tool functions, each corresponding to a different multimodal task:

  • describe_image: Used for image understanding. Supports screenshots, OCR, and chart analysis, and can process multiple images simultaneously. Uses the qwen3-vl-flash model by default.
  • transcribe_audio: Used for audio transcription. Supports common formats such as wav/mp3/m4a/aac/flac/ogg/amr. Uses the qwen3-asr-flash model by default.
  • generate_image: Used for generating images from text and saving them locally. Uses the qwen-image-plus model by default.

During processing, all media content is kept out of the main model’s context. It is converted to text and returned, preserving the efficiency of the main model.

Installation and Configuration

Installation

Run the following command in the terminal to add the plugin to the specified profile:

dsh plugin --profile demo add github:wuwangmao/dsh-qwen-multimodal

If installing from a local directory or tarball, you can use the following commands:

dsh plugin --profile demo add ./dsh-qwen-multimodal
# or
pnpm pack
dsh plugin --profile demo add ./dsh-qwen-multimodal-0.1.0.tgz

Prerequisites

  1. Python Environment: Python 3.10 or higher must be installed (the image generation script uses int | None type annotations).
  2. API Key: You need to configure the Alibaba Cloud Bailian API Key. The plugin reuses this key for vision, speech, and image generation.
  3. Python Path: The plugin resolves the python command from the system PATH by default. If resolution fails, override the relevant line in the profile’s cordis.patch.yml and set config.pythonPath.

Configuration File

After installation, configure the environment variable file. Copy the example file and edit it:

cp skills/deepseek-vision/.env.example skills/deepseek-vision/.env

Add VISION_API_KEY to the .env file (create it in the Alibaba Cloud Bailian console).

Usage Examples

After the plugin is loaded, the model can directly call the three tools above. The following are typical usages:

  1. Image Understanding and OCR
    describe_image({ images: ['screenshot.png'] })
    // Extract text, code, or error information from the image
    describe_image({ images: ['chart.png'], prompt: '逐字提取图中所有文字,保留原样' })
    // Use a custom prompt to obtain more precise output
  1. Speech-to-Text
    transcribe_audio({ audios: ['recording.m4a'], language: 'zh' })
    // Specify the language to improve transcription accuracy
  1. Text-to-Image Generation
    generate_image({ prompt: 'a cute orange cat on a windowsill watching the sunset', out_dir: './out' })
    // Generate an image and save it to the specified local directory
  1. Generation Quality Check
    generate_image({ prompt: '...', out_dir: './out', verify: true })
    // Automatically use Qwen VL after generation to perform a quality check and ensure the result matches the prompt requirements

Notes

  • Permissions and Dependencies: The plugin runs under the permissions of the current DSH process. Review the source code before installation.
  • Environment Variable Location: It is recommended to place the .env file in an external directory (for example, D:/qwen-vision) and point to it via skillDir in cordis.patch.yml, so that secrets are not stored in node_modules.
  • License: This project is licensed under the MIT License.

Conclusion

wuwangmao/dsh-qwen-multimodal solves the pain point of quickly integrating multimodal capabilities in DSH. By encapsulating Qwen’s API capabilities as standard DSH tools, developers can focus on model logic without handling low-level API calls and file I/O details.

Plugin directory: https://www.skillhub.cn/plugins/wuwangmao/dsh-qwen-multimodal
GitHub repository: https://github.com/wuwangmao/dsh-qwen-multimodal