DeepSeek Harness Plugins · Model Inference
Plugins that extend DeepSeek Harness (DSH) with vision, workflows, memory, web tools and more.
An auto-discovery CLIProxyAPI provider for DeepSeek Harness that syncs live model catalogs and handles vision preprocessing, simplifying model selection.
dsh-llm-opencode is an OpenCode Zen free model adapter for DeepSeek Harness, allowing coding agents to use free models without an API key, ready to use out of the box.
dsh-plugin-rollout-scout is a rollout scout plugin that analyzes streaming chain-of-thought output and paragraph openings, using heuristic scoring based on phrasing to identify new chat models in gray-release deployments.
meow-vision offers image recognition for text models and UI preview screenshots for multimodal models, supporting component design iteration.
An LLM adapter plugin for DeepSeek Harness that accesses WorkBuddy/CodeBuddy models via a local proxy, requiring no API key, and includes a web status widget.
A dsh plugin implementing Best-of-N self-verification to enhance LLM inference quality like frontier models at a fraction of the cost.
This plugin provides cancelable auxiliary model calls, durable limits, and provenance-preserving usage for DeepSeek Harness.
A community DeepSeek Harness plugin that exposes DSH Agents via the AG-UI protocol, providing an authenticated HTTP/SSE gateway, thread-to-agent bindings, streamed text and tool events, browser-owned tools, and turn continuation.
Provides a visual interface for DeepSeek Harness to inspect the details and context sources of each AgentLoop model call request.
This plugin provides full-coverage image tiling for DeepSeek Harness vision models, enhancing recognition of small text in high-resolution images.
dsh-model-failover is a plugin for DeepSeek Harness that implements a two-level model circuit breaker with failover, automatically routing requests to fallback models when failures occur, without requiring core modifications.
dsh-whale is a slim mode plugin for DeepSeek Harness that trims verbosity from AI responses, offering multiple compression levels to save 60-75% of output tokens while preserving all technical information.