AI Agent Hub
Back to plugins
🤖

dsh-context-budget

Model Inference Updated 2026.09.09

Run the following command in DeepSeek Harness:

dsh plugin install zpis666/dsh-context-budget

Paste the following prompt into your AI chat to install this plugin:

Run dsh plugin install zpis666/dsh-context-budget in the DeepSeek Harness terminal to install the plugin; the source code is hosted at https://github.com/zpis666/dsh-context-budget .

About this plugin

Hand-declared third-party gateway routes rarely match the built-in pi-ai catalog, so the context window silently falls back to the 262144 default; meanwhile the compaction backend only accepts a ratio, leaving no direct way to express a compact-at-600k trigger. dsh-context-budget consolidates these scattered settings into a single field group on the Settings → Models provider card: context window, compaction threshold, retain budget, and thinking effort. Fill them once and the values hot-apply in about a second with no second restart.

Core capabilities: the plugin converts your absolute threshold into the ratio the backend expects and writes it to cordis.patch.yml; it places a thinking-effort picker at the bottom-right of the composer, auto-matching vendor profiles for xAI, OpenAI, Anthropic, DeepSeek, Google, and others; every write is validated for correctness, preserves all unrelated lines in the file, and rolls back immediately on failure with a backup left behind.

Ideal for developers running pi-ai routes through self-hosted or third-party gateways who need precise control over context capacity and compaction triggers, and who want to manually set reasoning-effort levels for models from different vendors.

Use Cases

  • Set exact context window and compaction trigger for self-hosted gateway routes
  • Pick reasoning-effort levels per vendor without digging through settings
  • Fill capacity fields once on the provider card and see them apply live

Best For

  • Developers running pi-ai routes through third-party or self-hosted gateways
  • Inference users who need precise control over context capacity and compaction
  • Ops staff adjusting model parameters without a service restart