Introduction

The design philosophy of DeepSeek Harness (DSH) is “everything is a plugin,” allowing system functionality to be extended through plugins. For developers, local or open-source chat models usually accept text input only (inputModalities is text). When the native model does not support image or other file formats, DSH’s default behavior is to directly reject image input (error Model does not support image input) or to make it impossible to attach videos, audio, and ordinary files.

The dsh-multimodal plugin solves this problem. It intercepts input before the file is passed to the chat model, using a preset model chain to convert images, videos, or audio into textual descriptions (Prompt Tokens), which are then processed by the chat model. This allows users to leverage vision models or other processing models to preprocess multimodal content.

Core Features

This plugin adds a standalone Multimodal page to DSH settings, providing the following capabilities:

  1. Independently configured processing chains: Configure separate processing model chains for different file types (images, videos, audio, and text). Each preset can contain multiple models, attempted in order, with the first successful one winning.
  2. Preset management: The plugin includes a built-in default preset “Multimedia (Default)” that covers common image (png/jpg/webp/gif), video (mp4/webm, etc.), and audio formats. Users can create multiple presets and match specific file types using wildcards.
  3. Native text pass-through: There are no presets for the text type; file content is always passed as-is without conversion.
  4. Smart fallback: If the configured model chain is empty or processing fails, the system automatically falls back to the current session’s default model.
  5. Prompt configuration: The processing prompt for a preset can be customized. If left blank, the system uses the default prompt defined by the system for that file type.
  6. Model capability validation: The processing model must declare image capability. The Provider adapter decides whether to accept images based on the model’s declared input modalities; otherwise, it directly rejects them (for example, when a local model is not configured with input: [text, image]).
  7. Processing switch and paths:
    • No processing by default: Images are not automatically processed by default, fully following native behavior (native multimodal models pass through unchanged, text-only models reject them).
    • Two paths: Supports the “send path” (pasting/dragging an image) and the “attachment path” (the ➕ button in the input box).
  8. Size limit: The default size limit is 10 MB, and each preset can override it individually.
  9. Failure policy: Each preset can use “pass-through” or “replace with a failure message”.

Installation

Use the official DSH CLI to install the plugin. Replace <本插件路径或 tgz> with the actual file path or Tgz package name.

# 在 web profile 中挂载
dsh plugin --profile web add <本插件路径或 tgz>

# headless profile 同样适用
dsh plugin --profile headless add <本插件路径或 tgz>

After installation, hard-refresh the browser (Ctrl+Shift+R) to see the “Multimodal” page in settings. If you modify host-side code, you may need to restart DSH.

Usage

After enabling and configuring the plugin, follow these steps:

  1. Open Settings → Multimodal.
  2. Turn on the “Enable Multimodal Processing” switch.
  3. Configure the model chain for a preset. In settings, specify the Provider and model name(s) (multiple lines are allowed); the order is the fallback order. Use the “Test” button to verify connectivity.
  4. Return to the chat interface:
    • Send path: Paste or drag an image directly. The image is immediately logged and converted into a text description during the first model call.
    • Attachment path: Click the ➕ button in the input box and select any file (video/audio, etc.); the system will process it automatically and store it in the draft.

Configuration Example

The configuration file is usually located in the DSH settings directory (for example, settings.yaml). The following is an example configuring an image preset, using deepseek-official and openai as fallback vision models:

plugins:
  multimodal:
    enabled: true
    presets:
      - id: image
        name: 图片
        kind: image
        patterns: []          # 空 = 匹配该类型所有文件
        prompt: 请详细描述这张图片…
        maxBytes: 10485760    # 覆盖大小上限(字节,默认 10MB)
        onError: note         # 默认替换为失败说明;或 pass-through
        models:
          - provider: deepseek-official
            model: deepseek-chat
          - provider: openai
            model: gpt-4o

Limitations and Notes

Before using it, be aware of the following limitations:

  1. Video/audio processing: The adapter layer currently only supports extracting text and image content blocks (i.e., metadata briefs containing file name/type/size/MIME). It cannot directly transcribe audio or video content. Transcription requires using appropriate tools in the session.
  2. Image formats: Only png / jpeg / webp / gif formats accepted by the DSH attachment service are supported.
  3. Remote browsers: The /multimodal channel is loopback-authorized and does not take effect for remote browsers.
  4. Pass-through policy: If processing fails and the policy is set to pass-through, the system restores native DSH behavior (for example, a text-only model rejects images).
  5. Default behavior: If no preset is configured or “Automatically process images when sending” is not enabled, images are processed according to the chat model’s capabilities; text/video/audio are not attached.

Conclusion

dsh-multimodal provides DSH with flexible multimodal preprocessing capabilities. By inserting a vision model chain before the chat model, it solves the issue where native text models cannot process image attachments. Developers can configure different presets and fallback strategies as needed to accommodate various locally deployed model combinations.