In DeepSeek Harness development, how to switch between “economy models” and “high-quality models” based on task complexity is a common requirement. Manually implementing classification logic increases system complexity and makes it difficult to ensure consistency. The dsh-adaptive-model-router plugin provides a deterministic, localized routing solution that can be enabled through simple configuration to switch models on demand.
What Is This¶
dsh-adaptive-model-router is a Deterministic per-turn adaptive model routing plugin maintained by icyaaaww. It is designed for DeepSeek Harness 0.1.1-rc.2.
The core value of this plugin is that it does not classify requests by calling another model; instead, it uses local, synchronous, reproducible computation to determine routing. It automatically upgrades the model based on task characteristics within a single session, but does not downgrade it, thereby balancing cost control and response quality.
Core Features¶
- Deterministic per-turn routing: At the start of each turn, determines whether to switch to a high-quality model based on input character count and keywords.
- Local synchronous execution: The routing logic is fully localized and does not make additional model calls, so it does not increase token consumption or generate extra provider requests.
- Hysteresis behavior: Within a single session (turn), routing can upgrade but not downgrade. Re-evaluation and downgrading to an economy model can only occur when entering a new conversational turn.
- Flexible upgrade conditions: Supports triggering upgrades based on step count (
upgradeAfterStep) or tool failure count (upgradeAfterToolFailures). - Parameter control: Supports explicitly setting
reasoningEffortandmaxTokensin routing and allows preserving unconfigured routing selections throughpreserveUnknownSelection. - Transparency: The plugin records changes to request configuration and maintains consistency of
providerandmodelvariables, making debugging and replay easier.
Installation and Enablement¶
Installing this plugin requires using the DeepSeek Harness plugin installation command. Please ensure your environment meets Node.js ^22.19.0 || >=24.0.0.
dsh plugin --profile web add github:icyaaaww/dsh-adaptive-model-router
Configuration and Typical Usage¶
After installation, you need to define the economy model and high-quality model in the Harness configuration and set the routing rules.
Basic Configuration Example¶
Define the router in the configuration file, set the provider and model for economy mode and quality mode, and specify the threshold that triggers an upgrade.
- id: adaptive-model-router
name: dsh-adaptive-model-router
config:
economy:
provider: deepseek-official
model: deepseek-v4-flash
quality:
provider: deepseek-official
model: deepseek-v4-pro
inputCharsThreshold: 1200
complexityKeywords: [architecture, migration, security, 架构, 迁移, 安全]
upgradeAfterStep: 2
upgradeAfterToolFailures: 1
failureExclude: [todo_write, job_output, job_list]
preserveUnknownSelection: true
Configuration Item Description¶
- economy / quality: Defines the provider and specific model name for the low-cost model (such as flash) and the high-cost model (such as pro).
- inputCharsThreshold: Triggers an upgrade when the input character count reaches this value.
- complexityKeywords: Triggers an upgrade when these keywords are included (case-insensitive).
- upgradeAfterStep: Counts from 0 and triggers an upgrade when the step count reaches this value.
- upgradeAfterToolFailures: Triggers an upgrade when the number of consecutive non-excluded tool failures reaches this count.
- failureExclude: Specifies the list of tools to exclude when counting tool failures.
- preserveUnknownSelection: If set to
true, requests that do not match a routing rule remain unchanged; if set tofalse, the plugin takes over all requests.
Parameter Adjustment¶
The routing configuration supports optional parameters. Specifying reasoningEffort in routing clears the previous model’s setting; not specifying maxTokens preserves the explicit request limit.
Notes¶
- Complexity estimation: Keyword and length rules are only used to estimate complexity and cannot directly measure answer quality.
- Tool failure counting: Unless excluded in
failureExclude, tool failures caused by policy rejection are also counted toward the upgrade conditions. - Parallel tools: If routing upgrades in the middle of tool execution, tools that are already running in parallel may still be using the previous model.
- Model catalog validation: The plugin does not validate whether a model exists in the catalog; the selected provider is responsible for availability diagnostics.
- Cost reporting: The plugin does not include cost reporting functionality; effect should be evaluated by comparing telemetry data before and after deployment.
Applicable Scenarios¶
This plugin is suitable for scenarios that require dynamically adjusting model complexity during a conversation. For example, simple question-answering can use the flash model to reduce costs, while complex architectural design or tool call failures can automatically trigger a switch to the pro model, thereby controlling inference costs while ensuring final quality.