Introduction

The context compression feature in DeepSeek Harness can prevent conversation history from exceeding the model’s context window, but the default configuration requires editing a YAML file, which is not intuitive enough. deepseek-harness-compaction-ui is a visual configuration plugin that encapsulates the configuration items that would otherwise require manual editing into a Web settings page. It supports both percentage and absolute Tokens modes, making it easier to adjust the automatic compaction threshold and the amount of recent content to retain.

Core Features

  1. Visual Configuration: Manage context compression in “Settings → Plugins → Plugin Configuration” instead of manually editing YAML.
  2. Dual-Mode Support:
    • Trigger Threshold: Supports triggering compression based on a fixed number of Tokens or a percentage of the current context window.
    • Recent Content Retention: Supports retaining content based on a fixed number of Tokens or a percentage of the current context window.
  3. Automatic Conversion: When switching between Tokens and percentage modes, the equivalent value is automatically calculated and displayed.
  4. Hot Update: After saving the configuration, no need to recreate the session; the policy takes effect immediately.
  5. Runtime Monitoring: Displays a compression threshold progress ring next to the input box. Hovering over it shows the current context usage, automatic compaction threshold, and occupancy percentage. The ring turns red when the threshold is reached.
  6. Independent Policies: Supports setting independent compaction policies for different providers and models.

Installation and Activation

Use the following command to install the plugin into the DSH Web profile:

npx @deepseek-ai/dsh plugin --profile web add deepseek-harness-compaction-ui

After installation, open “Settings → Plugins → Plugin Configuration → Context Compression” in the Web interface to use it.

Typical Usage

After entering the plugin configuration page, select “Tokens Mode” or “Percentage Mode” as needed.

  • Tokens Mode: Directly set the absolute Token counts for triggering and retention. The percentage mode also displays an approximate Token count converted based on the current context window.
  • Percentage Mode: Sets ratios based on the model’s context window.
  • Configuration Logic: The configured “recent conversation retention” amount must be less than the “trigger threshold”; otherwise, it cannot be saved.

The default policy is as follows:

Setting Default Value
Automatic Compaction Enabled
Trigger Threshold 256,000 Tokens
Recent Conversation Retention 64,000 Tokens
Context Window 1,000,000 Tokens

Implementation Boundaries

This plugin is not a standalone compaction engine. It does not handle logic for selecting the compaction scope, tool call pairing, summary generation, message replacement, persistent events, cancellation, or context overflow recovery. Its core responsibilities are:

  1. Display and save visual settings.
  2. Convert absolute values into the thresholdRatio used by the official engine.
  3. Pass percentage-based policies directly to the official engine.
  4. Validate configuration and hot-update the policy at runtime.

Notes

Context compression is mainly used to control the growth of historical messages and cannot solve all capacity issues. The following scenarios may still lead to overflow or errors:

  • A single message or tool result itself exceeds the budget.
  • The System prompt and tool Schema are already close to the context limit.
  • The summarization model keeps failing or an external service is unavailable.
  • Information loss caused by multi-round summarization.

In addition, critical task state should be written to repository files or other persistent storage, and should not be saved only in conversation history.

Advanced Configuration

Manual YAML editing is usually not required. If you need to configure independent windows or policies for custom models, you can override the plugin configuration in the profile’s cordis.patch.yml.

- id: compaction-absolute-tokens
  config:
    thresholdMode: ratio
    thresholdRatio: 0.8
    retainMode: ratio
    retainRatio: 0.16
    thresholdTokens: 256000
    retainTokens: 64000
    contextWindowTokens: 1000000
    maxTokens: 8192
    compactionRetries: 1
    maxOverflowRetries: 1
    auto: true
    targets:
      - provider: deepseek-official
        model: deepseek-v4-flash
      - provider: deepseek-official
        model: deepseek-v4-pro

References