Introduction

The core design philosophy of DeepSeek Harness (DSH) is “everything is a plugin.” In multi-model orchestration scenarios, how to automatically select the optimal model tier based on task type and complexity is a typical engineering problem. The dsh-agentic-router plugin is designed to solve this problem. It introduces multi-expert recommendation and meta-selection mechanisms, combined with a reward system based on quality signals, to achieve automatic learning and optimization of routing strategies.

Installation

Installing the plugin requires specifying a DSH profile.

dsh plugin --profile web add dsh-agentic-router

After installation, the Web service must be restarted for the tool schemas to enter the prompt assembly process.

Core Features

Routing Architecture

The plugin uses a four-expert parallel-recommendation plus EXP3 meta-selector architecture:
* Four experts: rule (rule prior), UCB (multi-armed bandit), LinUCB (linear context), and k-NN (nearest neighbor).
* EXP3 selector: makes the final decision at the start of a turn based on the recommendations of the four experts.
* Data flywheel: At the end of each turn, a reward is calculated and fed back based on quality-proxy signals (tool failures, model retries, cost, latency), which updates the policy weights.

Real Model Switching (v1.4.0+)

v1.4.0 fixed the defect in previous versions that prevented real model switching. Now, by appending a request/header event (reason: router) to the session log during the agent/inbox/claimed stage, combined with DSH’s model selection priority chain (current-process selection → latest header in session log → default model), real model switching is enabled.

Step-level Routing (v1.5.0)

v1.5.0 introduces step-level routing. During the agent/pre-step stage, the plugin checks the tool type invoked in the previous step:
* Heavy tools (e.g., code writing, execution, sub-agent): upgraded to the strong tier in the next step.
* Lightweight/text-only tools: fall back to the baseline tier set at the turn level.

This adjustment applies only to the next step, and the data is persisted to steps.jsonl.

Toolset

The plugin provides the following tools for control and observation:
* agentic_router_stats: View decision records, four-expert recommendations, EXP3 weights, and flywheel status.
* agentic_router_set_mode: Switch mode (shadow/active/off).
* agentic_router_force: Force a specific model ID (effective in active mode).
* agentic_router_reset: Clear the flywheel learning state.
* agentic_router_set_prices: Configure model prices.

Configuration and Usage

Mode Switching

The default is shadow mode (logging and learning only, no model switching). Switch to active mode (real switching) using the following methods:
1. In-session switch: In any DSH session, input “Switch agentic router to active mode.” The plugin calls agentic_router_set_mode and persists the setting to ~/.dsh/storages/dsh-agentic-router/policy.json.
2. Configuration file switch: Configure it in the profile’s cordis.patch.yml:

- id: agentic-router
  config:
    mode: active

Tier Recognition

The model tier is recognized by keywords in the model ID:
* flash / fast: routed to the fast tier.
* pro / strong: routed to the strong tier.
* Otherwise: defaults to the mid tier.

Data Flywheel

Data Persistence

All decisions and rewards are persisted in the ${DSH_HOME:-~/.dsh}/storages/dsh-agentic-router/ directory:
* decisions.jsonl: Records the task type, complexity, expert recommendations, meta-layer selection, and actual routing for each decision.
* rewards.jsonl: Records reward details (tool failures/retries/latency) and feature vectors. After restart, replay restores the learning state.
* policy.json: Stores mode and forced-model configurations.

Feedback Mechanism

  • Settlement window: The feedback settlement window is 45s (configurable via settleDelayMs).
  • Reward calculation: Starts at 1 point for a clean completion; tool failure -0.2/time (capped at 0.6); model retry -0.2/time (capped at 0.4); cost min(0.5, 20 × turn cost in yuan); latency >30s -0.1 / >90s -0.3.
  • Human feedback: Relies on the 👍/👎 buttons in the Web UI (👍 +0.1, 👎 -0.5), used to adjust rewards that have already been applied.

Notes

  • Restart loss: Unsettled pending feedback entries are lost after a restart.
  • Feedback dependency: Feedback data depends on the deployed Web UI actually generating it. The plugin only reads and does not write.
  • Expert graduation: LinUCB and k-NN experts require a certain number of turn samples (50/30) before entering the meta-layer pool.
  • Required parameter: The install command requires --profile.

References