Introduction¶
DeepSeek Harness (DSH) focuses its reasoning capabilities on DeepSeek models. When visual tasks need to be handled, manually configuring providers is both cumbersome and error-prone. The hanchn/dsh-multimodal-router plugin solves this problem by delegating image perception to local or explicitly configured multimodal models and returning textual evidence to the reasoning Agent. It follows a “local-first” philosophy and disables remote fallback by default to protect privacy.
What It Is¶
This is a plugin maintained by hanchn under the MIT license. It is a zero-configuration tool that automatically discovers vision-capable models using Ollama’s model metadata and supports Ollama and OpenAI-compatible vision APIs. The plugin decouples visual processing from DeepSeek reasoning through its routing capabilities, allowing the Agent to focus on analyzing textual evidence.
Core Features¶
- Automatic discovery and routing: Automatically discovers vision-capable models via Ollama’s model metadata, supports local-first routing, and disables remote fallback by default.
- Multimodal interaction: Provides an image attachment button, supporting drag-and-drop and clipboard pasting (PNG, JPEG, WebP, GIF, image data URL).
- Voice conversation capability: Model-capability-based voice conversation, supporting audio processing for Gemma 4 E2B, E4B, and 12B; raw audio is sent only to local Ollama and is not archived.
- Presets and tools: Built-in presets for general description, OCR, UI, charts, visible code/errors, and an explicit
inspect_imagetool. - Privacy protection: API keys are read from environment variables and are not stored in plugin configuration; raw audio is not uploaded to the cloud.
Installation and Activation¶
Before installation, ensure that Node.js 22.19+ is installed on the system and that DeepSeek Harness version 0.1.0-rc.6 is running.
- In the plugin directory, copy the environment variable template and configure the key:
cp .env.example .env
# Edit .env locally and set DEEPSEEK_API_KEY. Never commit .env.
- Add the plugin to DSH’s Web configuration:
npx @deepseek-ai/dsh@0.1.0-rc.6 plugin --profile web add .
- Start the DeepSeek Harness Web service:
npx @deepseek-ai/dsh@0.1.0-rc.6 web
After startup, open http://127.0.0.1:3080 to start using it.
Typical Usage¶
- UI interaction: Use the paperclip button on the page to select an image, drag an image into the input box, or paste directly with
Cmd/Ctrl+V. The plugin automatically analyzes the attachment, and the DeepSeek model answers based on the analysis results. - Command-line tool: Use the
inspect_imagetool to analyze a local file path. For example:
Use inspect_image to analyze /absolute/path/to/screenshot.png, then explain the UI issue.
- Voice settings: The plugin uses the
Serenavoice by default. To use another built-in voice, settts.voicetoVivian,Dylan,Uncle_Fu, orEricin the configuration.
Configuration and Advanced Settings¶
The plugin supports configuring multiple priority providers. By editing the plugin entry in DSH’s profile patch, you can customize Ollama and OpenAI-compatible vision API providers.
- id: multimodal-router
name: '@hanchn/dsh-multimodal-router'
config:
automaticAttachments: true
archiveDirectory: .dsh-multimodal-router/images
cacheDirectory: .dsh-multimodal-router/cache
discoveryCacheMs: 300000
resultCacheMs: 3600000
resultCacheMaxEntries: 64
ollamaKeepAlive: 2m
ollamaIdleUnloadMs: 60000
maxVisionTokens: 256
providers:
- id: local-auto
type: ollama
baseUrl: http://127.0.0.1:11434
model: auto
apiKeyEnv: ''
priority: 100
- id: company-vlm
type: openai-compatible
baseUrl: https://vision.example.com/v1
model: qwen-vl
apiKeyEnv: COMPANY_VLM_KEY
priority: 50
allowRemoteFallback: false
timeoutMs: 180000
maxImageBytes: 20971520
realtimeAudio: true
transcriptionLocale: zh-CN
audioChunkMs: 1500
maxAudioSeconds: 30
maxAudioBytes: 10485760
tts:
enabled: true
autoStart: true
baseUrl: http://127.0.0.1:8000
model: mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit
voice: Serena
responseFormat: wav
browserFallback: false
maxTextCharacters: 8000
keepAliveMs: 300000
Configuration notes:
- allowRemoteFallback: true: Enable only when it is confirmed that images may automatically leave the machine; defaults to false.
- automaticAttachments: true: Exposes the routed vision capability to DSH’s admission layer and enables automatic attachment processing; set to false to use only the explicit tool.
- archiveDirectory: Stores content-addressed local copies of automatically processed images; the default directory is ignored by Git by default.
Applicable Scenarios and Notes¶
- Environment requirements: Node.js 22.19+ and DeepSeek Harness
0.1.0-rc.6are required. The default path requires Ollama to be running a multimodal model on127.0.0.1:11434. - Voice limits: Gemma 4 has a 30-second audio limit per utterance. Browser/OS
speechSynthesisfallback is disabled by default. - Optional voice service: The optional high-quality voice reply feature requires MLX-Audio to be running on
127.0.0.1:8000.
Conclusion¶
hanchn/dsh-multimodal-router provides lightweight vision and audio processing capabilities for DeepSeek Harness through zero-configuration local discovery and local-first routing. It is suitable for developers who need to protect privacy in local environments and do not want to manually manage vision model providers.