In DeepSeek Harness (DSH), every turn of the conversation resends the entire history. As the conversation grows longer, input token costs can rise quickly and become the dominant overhead. dsh-squeeze-command is designed to address this issue. It provides manual, budget-oriented context compression: the dialogue model selects the scope, and a cheap, fast model (flash-class) writes the summary.

What It Is

This is a DSH plugin for compressing session context on demand. It is maintained by hardes11. It reduces input token costs by allowing the dialogue model to select the scope and having a cheap, fast model write the summary, while retaining the original session log.

Core Features

  • Cut resent content: Squeeze the conversation into the budget range with a single command, and generate the summary using the configured cheap routing rather than the expensive dialogue model.
  • No dropped content: Each compression leaves a checkpoint on the surface, while original messages remain in the append-only session log; use /squeeze map to browse them.
  • Resist post-compression confusion: Each checkpoint summary includes explicit status lines and calculated reliable summaries from uncompressed tails, preventing the model from re-entering work that is described in the summary but not actually completed.
  • Triggers only when you ask: There are no thresholds, no automatic triggers, and it refuses to run while an agent turn is in progress.
  • No agent scheduling required: Summary delegation is built in; you only need to run a command, and everything else happens within it.

Install and Enable

  1. Clone the plugin into the DSH commands directory:
git clone https://github.com/hardes11/dsh-squeeze-command.git ~/.dsh/commands/dsh-squeeze-command
  1. Mount and configure it in the preset agent.cordis.yml. You need to specify the target token budget and summary model:
- id: command-squeeze
  name: 'dsh-squeeze-command'
  config:
    contextBudgetTokens: 100000
    summarizerProvider: <your-provider>
    summarizerModel: <your-flash-model>

Command Usage

  1. /squeeze: Compress to the configured contextBudgetTokens (preset configuration).
  2. /squeeze 60k: Specify a one-time compression target (for example, 60000 tokens).
  3. /squeeze status: View the current session’s token estimate, budget, and summarizer status.
  4. /squeeze map: Generate an interactive HTML map to compare before and after compression and view retention ratio.
  5. /squeeze help: View the full usage guide.

Notes

  • Manual mode only: The plugin refuses to run while an agent turn is in progress.
  • Cache invalidation: Compression rewrites the context prefix, invalidating the prompt cache; it is recommended to run it when the cache has expired or caching is not needed.
  • Model requirements: The summary routing should be configured as a cheap, fast model (flash-class recommended), because summary tasks do not require high reasoning intensity.
  • Required configuration: Configure contextBudgetTokens, summarizerProvider, and summarizerModel before running.

Summary

dsh-squeeze-command is a practical tool in DeepSeek Harness for controlling context costs. It reduces token consumption while maintaining context integrity through manual triggering and a flash-class summary model, making it suitable for long-running, input-cost-sensitive session scenarios.