dsh-iris
Run the following command in DeepSeek Harness:
dsh plugin install mokuyoaxis/dsh-iris
Paste the following prompt into your AI chat to install this plugin:
Run the install command in your DeepSeek Harness environment to load the plugin; the source repository is available at https://github.com/mokuyoaxis/dsh-iris.
About this plugin
Iris consolidates image generation, text-to-video, image-to-video, S2V digital-human video, text-to-speech, audio transcription, long-image OCR, object localization, pixel-difference analysis, video frame extraction, and multimodal summarization into a single runtime, so an Agent no longer needs a separate vendor API integration per task. Under the hood it builds a multi-provider model pool, assigns candidate models per capability dimension, supports DashScope and OpenAI-compatible image services, and automatically fails over when submission fails. An async task system ensures that even after a plugin restart, remote tasks still in flight are picked up without duplicate billing.
Management lives in one Iris workbench: add providers, discover models, assign capabilities, track task status, and play media artifacts via authorized-token links. File input supports three sources (browser upload, session attachment, host path). Provider credentials stay on the host side with tightened file permissions, and the API never returns full keys. Security boundaries include restricting DashScope credential transmission to official Alibaba Cloud domains, requiring explicit trusted-host configuration beyond loopback, and rendering HTML screenshots in a sandbox with no remote scripts.
It is aimed at teams and individual developers already running Agent workflows on DeepSeek Harness who want to fold multimodal media production into a unified scheduling layer. No standalone server or extra port is needed; the plugin registers server-side tools and a Web client slot into the host, ready to use on demand.
Screenshots
Use Cases
- Agents invoke image, video, and speech generation directly from conversation without manual vendor API wiring
- Run long-image OCR, object localization, and pixel-diff comparison for automated visual QA
- Extract key frames from video and combine with audio transcription to produce multimodal summaries
Best For
- Teams running Agent workflows on DeepSeek Harness that need multimodal media capabilities
- Indie developers who want a single configuration layer over multiple vision and media providers
- Engineering teams requiring async task tracking and failover for production-grade reliability
Related Plugins
A unified suite combining hot runtime injection, task-aware thinking-mode routing, and a graded session protocol with red-team gates to sustain model diligence across long-horizon inference.
ModLens is a vision plugin for DeepSeek Harness that gives text-only models sight by reading images pasted directly into chat, with zero-config setup and multiple vision engines.
On-demand vision for text-only DeepSeek Harness agents: built-in free keyless vision chain and 14 vision tools, routing image turns as tool calls to vision models with pixel fidelity, no Python needed, one-command install.
Give text-only models in DeepSeek Harness eyes, enabling image Q&A, long-screenshot OCR, UI restoration, and GUI visual tasks.