Introduction

In multi-step agentic tasks built with DeepSeek Harness (DSH), models typically perform a “rethink” before each tool call. This mechanism can improve accuracy in complex tasks, but it also significantly increases the time cost (for example, a 50-step tool-call task may take several minutes due to intermediate reasoning). This plugin aims to finely control this process, preserving performance on complex tasks while avoiding wasted compute on simple tasks.

Plugin Overview

This plugin is an extension of DeepSeek Harness (DSH), maintained by the developer drscrewdriver. It mainly controls the reasoning_effort parameter in model requests, supports automatically scheduling the reasoning effort based on history, and allows manually pinning a specific effort level.

Core Features

The plugin provides the following main capabilities:

  1. Per-turn reasoning effort control: Adjust the model’s reasoning effort for a single conversation turn.
  2. Automatic scheduling mode: After selecting Auto in the session model selector, the plugin automatically schedules low, high, or max reasoning effort based on recent tool-call history.
  3. Manual locking mode: Supports manually fixing the reasoning effort, including off, on, minimal, low, medium, high, xhigh, and max.
  4. Version dependency: The plugin depends on DSH environment version >= 0.1.7-rc.1 < 0.1.8.

Usage

In the DSH session model selector, it can be configured in the following ways:

  1. Automatic scheduling:
    Select Auto in the model selector. The plugin then intervenes as a “mask,” parses the history, and decides whether to use low, high, or max for each step. This prevents simple tasks from becoming expensive while ensuring heavy tasks are not interrupted due to insufficient resources.

  2. Manual selection:
    Manually choose a specific level. For example, selecting high directly sends the official default reasoning effort value.

Supported reasoning levels:
* off: Disable thinking (manual mode only; it is not automatically selected).
* on: Enable thinking (only applicable to model-switching-only scenarios; sends the enable_thinking flag and does not send reasoning_effort).
* minimal: Minimal effort (suitable for lightweight tasks).
* low: Low effort (suitable for simple chat tasks).
* medium: Medium effort.
* high: High effort (official default).
* xhigh: Extra-high effort.
* max: Maximum effort (for heavy workloads).

Notes

  • Dependency change: Since v0.7.0-beta.1 (2026-09-06), the plugin no longer depends on the dsh-llm-openai-completions adapter and instead uses the official llm-pi-ai compatibility layer. If you are using an older DSH version (0.1.2–0.1.6), use plugin version 3.0.2.
  • Language support: Japanese (ja) and Korean (ko) are not available in the standard DSH environment. These languages require a custom DSH fork (updating locale-settings.ts and client/index.ts and rebuilding), or waiting for official support.
  • Level distinction: on is only used to enable the thinking toggle and is not a reasoning effort level. Do not select on for reasoning-capable models.

Summary

This plugin makes DSH agent execution more cost-effective and efficient by finely managing reasoning_effort. It is suitable for developers who need to optimize resource consumption while preserving reasoning quality on complex tasks.

Plugin directory
GitHub repository