In DeepSeek Harness development, how to switch between “economy models” and “high-quality models” based on task complexity is a common requirement. Manually implementing classification logic increases system complexity and makes it difficult to ensure consistency. The dsh-adaptive-model-router plugin provides a deterministic, localized routing solution that can be enabled through simple configuration to switch models on demand.

What Is This

dsh-adaptive-model-router is a Deterministic per-turn adaptive model routing plugin maintained by icyaaaww. It is designed for DeepSeek Harness 0.1.1-rc.2.

The core value of this plugin is that it does not classify requests by calling another model; instead, it uses local, synchronous, reproducible computation to determine routing. It automatically upgrades the model based on task characteristics within a single session, but does not downgrade it, thereby balancing cost control and response quality.

Core Features

  1. Deterministic per-turn routing: At the start of each turn, determines whether to switch to a high-quality model based on input character count and keywords.
  2. Local synchronous execution: The routing logic is fully localized and does not make additional model calls, so it does not increase token consumption or generate extra provider requests.
  3. Hysteresis behavior: Within a single session (turn), routing can upgrade but not downgrade. Re-evaluation and downgrading to an economy model can only occur when entering a new conversational turn.
  4. Flexible upgrade conditions: Supports triggering upgrades based on step count (upgradeAfterStep) or tool failure count (upgradeAfterToolFailures).
  5. Parameter control: Supports explicitly setting reasoningEffort and maxTokens in routing and allows preserving unconfigured routing selections through preserveUnknownSelection.
  6. Transparency: The plugin records changes to request configuration and maintains consistency of provider and model variables, making debugging and replay easier.

Installation and Enablement

Installing this plugin requires using the DeepSeek Harness plugin installation command. Please ensure your environment meets Node.js ^22.19.0 || >=24.0.0.

dsh plugin --profile web add github:icyaaaww/dsh-adaptive-model-router

Configuration and Typical Usage

After installation, you need to define the economy model and high-quality model in the Harness configuration and set the routing rules.

Basic Configuration Example

Define the router in the configuration file, set the provider and model for economy mode and quality mode, and specify the threshold that triggers an upgrade.

- id: adaptive-model-router
  name: dsh-adaptive-model-router
  config:
    economy:
      provider: deepseek-official
      model: deepseek-v4-flash
    quality:
      provider: deepseek-official
      model: deepseek-v4-pro
    inputCharsThreshold: 1200
    complexityKeywords: [architecture, migration, security, 架构, 迁移, 安全]
    upgradeAfterStep: 2
    upgradeAfterToolFailures: 1
    failureExclude: [todo_write, job_output, job_list]
    preserveUnknownSelection: true

Configuration Item Description

  • economy / quality: Defines the provider and specific model name for the low-cost model (such as flash) and the high-cost model (such as pro).
  • inputCharsThreshold: Triggers an upgrade when the input character count reaches this value.
  • complexityKeywords: Triggers an upgrade when these keywords are included (case-insensitive).
  • upgradeAfterStep: Counts from 0 and triggers an upgrade when the step count reaches this value.
  • upgradeAfterToolFailures: Triggers an upgrade when the number of consecutive non-excluded tool failures reaches this count.
  • failureExclude: Specifies the list of tools to exclude when counting tool failures.
  • preserveUnknownSelection: If set to true, requests that do not match a routing rule remain unchanged; if set to false, the plugin takes over all requests.

Parameter Adjustment

The routing configuration supports optional parameters. Specifying reasoningEffort in routing clears the previous model’s setting; not specifying maxTokens preserves the explicit request limit.

Notes

  • Complexity estimation: Keyword and length rules are only used to estimate complexity and cannot directly measure answer quality.
  • Tool failure counting: Unless excluded in failureExclude, tool failures caused by policy rejection are also counted toward the upgrade conditions.
  • Parallel tools: If routing upgrades in the middle of tool execution, tools that are already running in parallel may still be using the previous model.
  • Model catalog validation: The plugin does not validate whether a model exists in the catalog; the selected provider is responsible for availability diagnostics.
  • Cost reporting: The plugin does not include cost reporting functionality; effect should be evaluated by comparing telemetry data before and after deployment.

Applicable Scenarios

This plugin is suitable for scenarios that require dynamically adjusting model complexity during a conversation. For example, simple question-answering can use the flash model to reduce costs, while complex architectural design or tool call failures can automatically trigger a switch to the pro model, thereby controlling inference costs while ensuring final quality.