Introduction

When handling long conversational contexts in DeepSeek Harness (DSH), maintaining token usage efficiency is a common requirement. The dsh-openai-server-compaction plugin implements Codex-style compaction logic on the server side, providing automated context cleanup for OpenAI Responses routes managed by the llm-pi-ai plugin. It helps preserve conversational coherence without manual intervention.

Plugin Scope

This is a native DeepSeek Harness plugin designed to add persistent server-side compaction functionality to OpenAI Responses API routes. It is not responsible for connection management or model definitions, but focuses on executing compaction strategies and restoring state.

Core Features

The plugin provides the following capabilities:
* Server-side compaction: Supports manual triggering and automatic triggering based on token pressure and overflow, following the Codex Remote Compaction V2 protocol.
* History retention: Preserves the most recent 64,000 tokens of actual user messages during compaction.
* State persistence: Writes the compacted state into the DSH history and generates a portable text checkpoint.
* Request handling: Normal requests are processed via POST /v1/responses; compaction requests are triggered by appending a specific instruction to the request body.

Installation

Run the following command in the DSH environment to install the plugin:

dsh plugin --profile web add dsh-openai-server-compaction

Configuration and Usage

1. Configure the Route

The plugin does not create routes; it only processes routes that have already been configured. Under Settings -> Models, find the provider managed by the llm-pi-ai plugin and create or edit a route. The route must explicitly specify api: openai-responses, and configure baseURL and apiKeyEnv.

Configuration example:

llm-pi-ai:
  providers:
    openai:
      displayName: OpenAI Responses
      api: openai-responses
      baseURL: https://api.openai.com/v1
      apiKeyEnv: OPENAI_API_KEY
      models:
        - id: gpt-5.6
          name: GPT-5.6
          contextWindow: 1050000
          maxTokens: 128000

2. Enable Compaction

Under Settings -> Plugins -> OpenAI Compaction, enable the previously configured route and set the triggering threshold (thresholdRatio).

Configuration example:

openai-server-compaction:
  routes:
    openai:
      enabled: true
      thresholdRatio: 0.7

You must restart DSH after modifying the configuration.

Notes

  • Configuration dependency: The route must explicitly set api: openai-responses; otherwise, the plugin will not process the request.
  • State file: Compaction state is stored in $DSH_HOME/openai-server-compaction/state.json (default path). This is part of the operational configuration; do not delete it casually.
  • Output limitation: The plugin does not send the max_output_tokens parameter when making a compaction request.
  • Compatibility: This plugin indirectly handles OpenAI Responses routes via llm-pi-ai. If the route is misconfigured (for example, missing credentials or an unsupported model), the system will report a clear error.

Conclusion

This plugin is suitable for DSH users who need deep integration with the OpenAI Responses API and want automatic context length management. Source code and detailed documentation are available on GitHub.