DeepSeek Harness (DSH) adopts a plugin-based architecture, allowing extensions to alter model behavior. Reasoning level (reasoning_effort) determines how many tokens the model consumes across the Low/High/Max tiers, directly affecting cost and thinking depth. Manual configuration is usually a static, session-level setting, making it difficult to dynamically adjust according to the complexity of each message. The dsh-adaptive-effort plugin resolves this with an automatic scoring mechanism.

Plugin Overview

  • Name: dsh-adaptive-effort
  • Maintainer: imkingjh999
  • Category: Model inference
  • License: MIT
  • Core value: In DSH, automatically selects reasoning_effort (Low/High/Max) based on the complexity of each user message. It integrates MiniMax complexity scoring, token accounting (estimated downgrade savings), reply metadata badges, and includes compatibility handling for forced-thinking models such as GLM-5.3.

Core Features

  1. Automatic scoring and dispatch
    By default, it uses the MiniMax model to score user messages for complexity (thinking disabled, tokens limited). The score determines whether the current turn injects the low, high, or max tier. If scoring fails or times out, it automatically falls back to heuristic mode (rule-based judgment on length, code blocks, and keywords).

  2. Token accounting and estimation
    The plugin records each step’s tier (level), effort, reason, and scoring source. Each step’s usage (input/output tokens) is recorded in the ledger with a delay. The reporting feature aggregates token consumption by tier, estimates the “downgrade savings” in tokens, and flags suspected misrouting due to underestimation (Low-tier output explosion) or overestimation.

  3. Reply metadata badges
    Each AI reply ends with a badge showing the model, actual thinking tier, output tokens, and step count. Clicking the badge smoothly scrolls to the start of that reply. Badge data is sourced from the alignment between session snapshots and trajectory views, without consuming additional RPC calls.

  4. GLM-5.3 forced-thinking clamping
    Forced-thinking models such as GLM-5.3 do not support the medium tier, and thinking.type: disabled causes an error. The plugin clamps a manually selected off tier to low and logs it to prevent API call failures.

  5. Heavy-payload tool upgrade
    When tool arguments exceed 3,200 characters (heavy payload), the plugin forces the tier to max, ensuring long-argument tasks have sufficient thinking resources.

Installation and Enablement

In a local development environment, first clone or place the plugin directory, then run the installation command:

dsh plugin add ./dsh-adaptive-effort

After installation, refresh the DSH page for changes to take effect. Automatic scheduling must be enabled manually through the model selector.

Usage

1. Enable automatic scheduling

Open the model selection dialog, then in the “Reasoning level” list below the model, select “Automatic”. Only then does the plugin intervene, dynamically injecting reasoning_effort based on MiniMax scoring or heuristic rules. Selecting Off/Low/High/Max remains native manual tiers, which the plugin does not alter.

2. View CLI reports

Use Node.js to run the reporting script and view the tier distribution and savings statistics for a session:

node lib/report.js            # 会话索引
node lib/report.js --last     # 最近会话报表
node lib/report.js --session <id>
node lib/report.js --all      # 全部合并

Configuration Options

Plugin configuration supports cordis.yml or the adaptive-effort settings namespace. Common configuration keys are:

Key Default Description
enabled true Master switch. Controls whether automatic scheduling is enabled.
level auto Plugin-level tier. auto follows the model selector; off/low/high/max forces a global tier.
mode minimax Scoring backend. minimax uses LLM scoring; heuristic uses a pure rule-based fallback.
minimaxApiKey '' MiniMax API key. If empty, it reads the environment variable MINIMAX_API_KEY.
minimaxBaseUrl https://api.minimaxi.com/v1 International endpoint. For the China endpoint, change it to https://api.minimax.chat/v1.
minimaxModel MiniMax-M3 Scoring model used.
scoreTimeoutMs 4000 Timeout for scoring requests. On timeout, triggers heuristic fallback.
allowToolUpgrade true Allow heavy-payload tool arguments to upgrade the tier to max.
clampForcedThinkingOff true For forced-thinking models such as GLM-5.3, clamp manual off to low.
ledgerDir '' Ledger directory. If empty, defaults to ~/.dsh/adaptive-effort.

Notes and Limitations

  • Ledger collection delay: Accounting occurs during the next agent/request stage. If a turn is interrupted (for example, stopped before the assistant reply), that turn may be missing a usage record.
  • Heuristic fallback is conservative: The pure rule-based heuristic fallback relies on length and keywords, so the strategy is conservative.
  • Reports are read-only: The CLI report only reads local JSONL files and does not consume the API again.
  • Tier compatibility: Forced-thinking models such as GLM-5.3 do not support the medium tier; the plugin already clamps against this.

Summary

dsh-adaptive-effort provides DSH users with dynamic inference capability similar to an effort router. By combining MiniMax complexity scoring with token accounting, it preserves deep thinking for hard problems (max tier) while saving costs on trivial tasks (low tier) through downgrading. Combined with reply badges and reports, it lets you directly verify the actual tier consumption for each reply.

Plugin Directory | GitHub Repository