Introduction¶
The context compression feature in DeepSeek Harness can prevent conversation history from exceeding the model’s context window, but the default configuration requires editing a YAML file, which is not intuitive enough. deepseek-harness-compaction-ui is a visual configuration plugin that encapsulates the configuration items that would otherwise require manual editing into a Web settings page. It supports both percentage and absolute Tokens modes, making it easier to adjust the automatic compaction threshold and the amount of recent content to retain.
Core Features¶
- Visual Configuration: Manage context compression in “Settings → Plugins → Plugin Configuration” instead of manually editing YAML.
- Dual-Mode Support:
- Trigger Threshold: Supports triggering compression based on a fixed number of Tokens or a percentage of the current context window.
- Recent Content Retention: Supports retaining content based on a fixed number of Tokens or a percentage of the current context window.
- Automatic Conversion: When switching between Tokens and percentage modes, the equivalent value is automatically calculated and displayed.
- Hot Update: After saving the configuration, no need to recreate the session; the policy takes effect immediately.
- Runtime Monitoring: Displays a compression threshold progress ring next to the input box. Hovering over it shows the current context usage, automatic compaction threshold, and occupancy percentage. The ring turns red when the threshold is reached.
- Independent Policies: Supports setting independent compaction policies for different providers and models.
Installation and Activation¶
Use the following command to install the plugin into the DSH Web profile:
npx @deepseek-ai/dsh plugin --profile web add deepseek-harness-compaction-ui
After installation, open “Settings → Plugins → Plugin Configuration → Context Compression” in the Web interface to use it.
Typical Usage¶
After entering the plugin configuration page, select “Tokens Mode” or “Percentage Mode” as needed.
- Tokens Mode: Directly set the absolute Token counts for triggering and retention. The percentage mode also displays an approximate Token count converted based on the current context window.
- Percentage Mode: Sets ratios based on the model’s context window.
- Configuration Logic: The configured “recent conversation retention” amount must be less than the “trigger threshold”; otherwise, it cannot be saved.
The default policy is as follows:
| Setting | Default Value |
|---|---|
| Automatic Compaction | Enabled |
| Trigger Threshold | 256,000 Tokens |
| Recent Conversation Retention | 64,000 Tokens |
| Context Window | 1,000,000 Tokens |
Implementation Boundaries¶
This plugin is not a standalone compaction engine. It does not handle logic for selecting the compaction scope, tool call pairing, summary generation, message replacement, persistent events, cancellation, or context overflow recovery. Its core responsibilities are:
- Display and save visual settings.
- Convert absolute values into the
thresholdRatioused by the official engine. - Pass percentage-based policies directly to the official engine.
- Validate configuration and hot-update the policy at runtime.
Notes¶
Context compression is mainly used to control the growth of historical messages and cannot solve all capacity issues. The following scenarios may still lead to overflow or errors:
- A single message or tool result itself exceeds the budget.
- The System prompt and tool Schema are already close to the context limit.
- The summarization model keeps failing or an external service is unavailable.
- Information loss caused by multi-round summarization.
In addition, critical task state should be written to repository files or other persistent storage, and should not be saved only in conversation history.
Advanced Configuration¶
Manual YAML editing is usually not required. If you need to configure independent windows or policies for custom models, you can override the plugin configuration in the profile’s cordis.patch.yml.
- id: compaction-absolute-tokens
config:
thresholdMode: ratio
thresholdRatio: 0.8
retainMode: ratio
retainRatio: 0.16
thresholdTokens: 256000
retainTokens: 64000
contextWindowTokens: 1000000
maxTokens: 8192
compactionRetries: 1
maxOverflowRetries: 1
auto: true
targets:
- provider: deepseek-official
model: deepseek-v4-flash
- provider: deepseek-official
model: deepseek-v4-pro