DeepSeek’s main model is usually text-only. When tasks involve screenshots, charts, voice memos, video clips, or PDFs, a pure text model cannot directly process this unstructured data. The dsh-gemini-multimodal plugin solves this problem by acting as its “eyes and ears,” delegating visual and auditory tasks to Google Gemini. This enables DeepSeek Harness to perform image, audio, and video understanding, transcription, and image generation, while keeping reasoning logic with the main model.

Introduction

This is dsh-gemini-multimodal, maintained by RealAlexandreAI. The philosophy of DeepSeek Harness is “everything is a plugin.” This plugin fills the gap in DeepSeek models’ perceptual capabilities by integrating Gemini’s multimodal abilities. It does not modify the main model; instead, it exists as a tool layer, allowing the agent to “see” and “hear” when needed.

Core Features

The plugin provides the following capabilities:

  • Multimodal analysis: Analyze images / audio / video / URLs (supporting OCR, charts, UI interface recognition, and speech-to-text).
  • Transcription: Obtain verbatim text records of audio or video.
  • Image generation: Generate PNG images from text prompts and save them locally.
  • Document understanding: Summarize PDF / Office / text files or answer specific questions about them.

Installation and Enabling

The installation command is:

dsh plugin --profile web add dsh-gemini-multimodal

After installation, you need to enable the plugin in the DeepSeek Harness configuration.

Configuration and Providers

The plugin supports two provider modes, which must be specified in the configuration. If not configured, it defaults to antigravity_cli mode (provided that agy is installed locally).

You need to obtain an API key from aistudio.google.com. This mode calls the Google API directly over HTTP.

- id: gemini-multimodal
  name: dsh-gemini-multimodal
  config:
    provider: gemini_api
    api_key: <your gemini key>

2. Using antigravity_cli (default)

Use the locally installed agy command-line tool with Google account authorization, without needing to manage an API key. After installing agy, run it once and log in.

- id: gemini-multimodal
  name: dsh-gemini-multimodal
  config:
    provider: antigravity_cli

Typical Usage

After enabling the plugin, you can use the following tools in prompts or tool calls:

  • media_understand: Analyze input media content, such as screenshots or audio files.
  • media_transcribe: Extract verbatim text from audio or video.
  • image_generate: Generate images based on text descriptions.
  • read_document: Read PDF or Office documents and answer questions about them.

Notes

  • Permissions and security: When using antigravity_cli, interactive permission grants are not possible in headless mode, so the plugin passes the --dangerously-skip-permissions parameter by default. This grants full tool access for that run. For stricter permission control, set skip_permissions: false and add allow rules in your local configuration.
  • Privacy: When using gemini_api, media content is sent directly to Google in base64 format; when using antigravity_cli, media is read by a local proxy and then sent. The plugin itself does not log anything.
  • Rate limits: Free-tier image generation has rate limits and may return 429 errors.

Summary

dsh-gemini-multimodal delegates multimodal processing to Gemini, enabling DeepSeek Harness to handle image, audio, and document tasks. In scenarios that require DeepSeek models for complex reasoning while also processing non-text inputs, this is a practical plugin solution.

Plugin directory
GitHub repository