Introduction

DeepSeek Harness (DSH) focuses its reasoning capabilities on DeepSeek models. When visual tasks need to be handled, manually configuring providers is both cumbersome and error-prone. The hanchn/dsh-multimodal-router plugin solves this problem by delegating image perception to local or explicitly configured multimodal models and returning textual evidence to the reasoning Agent. It follows a “local-first” philosophy and disables remote fallback by default to protect privacy.

What It Is

This is a plugin maintained by hanchn under the MIT license. It is a zero-configuration tool that automatically discovers vision-capable models using Ollama’s model metadata and supports Ollama and OpenAI-compatible vision APIs. The plugin decouples visual processing from DeepSeek reasoning through its routing capabilities, allowing the Agent to focus on analyzing textual evidence.

Core Features

  • Automatic discovery and routing: Automatically discovers vision-capable models via Ollama’s model metadata, supports local-first routing, and disables remote fallback by default.
  • Multimodal interaction: Provides an image attachment button, supporting drag-and-drop and clipboard pasting (PNG, JPEG, WebP, GIF, image data URL).
  • Voice conversation capability: Model-capability-based voice conversation, supporting audio processing for Gemma 4 E2B, E4B, and 12B; raw audio is sent only to local Ollama and is not archived.
  • Presets and tools: Built-in presets for general description, OCR, UI, charts, visible code/errors, and an explicit inspect_image tool.
  • Privacy protection: API keys are read from environment variables and are not stored in plugin configuration; raw audio is not uploaded to the cloud.

Installation and Activation

Before installation, ensure that Node.js 22.19+ is installed on the system and that DeepSeek Harness version 0.1.0-rc.6 is running.

  1. In the plugin directory, copy the environment variable template and configure the key:
cp .env.example .env
# Edit .env locally and set DEEPSEEK_API_KEY. Never commit .env.
  1. Add the plugin to DSH’s Web configuration:
npx @deepseek-ai/dsh@0.1.0-rc.6 plugin --profile web add .
  1. Start the DeepSeek Harness Web service:
npx @deepseek-ai/dsh@0.1.0-rc.6 web

After startup, open http://127.0.0.1:3080 to start using it.

Typical Usage

  • UI interaction: Use the paperclip button on the page to select an image, drag an image into the input box, or paste directly with Cmd/Ctrl+V. The plugin automatically analyzes the attachment, and the DeepSeek model answers based on the analysis results.
  • Command-line tool: Use the inspect_image tool to analyze a local file path. For example:
  Use inspect_image to analyze /absolute/path/to/screenshot.png, then explain the UI issue.
  • Voice settings: The plugin uses the Serena voice by default. To use another built-in voice, set tts.voice to Vivian, Dylan, Uncle_Fu, or Eric in the configuration.

Configuration and Advanced Settings

The plugin supports configuring multiple priority providers. By editing the plugin entry in DSH’s profile patch, you can customize Ollama and OpenAI-compatible vision API providers.

- id: multimodal-router
  name: '@hanchn/dsh-multimodal-router'
  config:
    automaticAttachments: true
    archiveDirectory: .dsh-multimodal-router/images
    cacheDirectory: .dsh-multimodal-router/cache
    discoveryCacheMs: 300000
    resultCacheMs: 3600000
    resultCacheMaxEntries: 64
    ollamaKeepAlive: 2m
    ollamaIdleUnloadMs: 60000
    maxVisionTokens: 256
    providers:
      - id: local-auto
        type: ollama
        baseUrl: http://127.0.0.1:11434
        model: auto
        apiKeyEnv: ''
        priority: 100
      - id: company-vlm
        type: openai-compatible
        baseUrl: https://vision.example.com/v1
        model: qwen-vl
        apiKeyEnv: COMPANY_VLM_KEY
        priority: 50
    allowRemoteFallback: false
    timeoutMs: 180000
    maxImageBytes: 20971520
    realtimeAudio: true
    transcriptionLocale: zh-CN
    audioChunkMs: 1500
    maxAudioSeconds: 30
    maxAudioBytes: 10485760
    tts:
      enabled: true
      autoStart: true
      baseUrl: http://127.0.0.1:8000
      model: mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit
      voice: Serena
      responseFormat: wav
      browserFallback: false
      maxTextCharacters: 8000
      keepAliveMs: 300000

Configuration notes:
- allowRemoteFallback: true: Enable only when it is confirmed that images may automatically leave the machine; defaults to false.
- automaticAttachments: true: Exposes the routed vision capability to DSH’s admission layer and enables automatic attachment processing; set to false to use only the explicit tool.
- archiveDirectory: Stores content-addressed local copies of automatically processed images; the default directory is ignored by Git by default.

Applicable Scenarios and Notes

  • Environment requirements: Node.js 22.19+ and DeepSeek Harness 0.1.0-rc.6 are required. The default path requires Ollama to be running a multimodal model on 127.0.0.1:11434.
  • Voice limits: Gemma 4 has a 30-second audio limit per utterance. Browser/OS speechSynthesis fallback is disabled by default.
  • Optional voice service: The optional high-quality voice reply feature requires MLX-Audio to be running on 127.0.0.1:8000.

Conclusion

hanchn/dsh-multimodal-router provides lightweight vision and audio processing capabilities for DeepSeek Harness through zero-configuration local discovery and local-first routing. It is suitable for developers who need to protect privacy in local environments and do not want to manually manage vision model providers.

Plugin Directory
GitHub Repository