Introduction¶
When DeepSeek Harness (DSH) handles local models, the context window declared by the model is often not practical. As context length increases, prefill speed degrades significantly, causing Time to First Token (TTFT) to rise substantially, and even causing the GPU to hang before the window is fully filled. A 27B model may maintain 300 tokens/s before 100K context, but once it exceeds 100K, speed may drop to 70 tokens/s, causing subsequent replies to take minutes before they begin.
To address this issue, the dsh-context-budget plugin helps keep local model context within a range the GPU can handle well by measuring prefill speed, setting hard ceilings, and compressing early when limits are exceeded.
Core Features¶
This plugin is responsible for checking protected provider routes before each agent step to ensure context does not exceed GPU processing capabilities.
- Context control: Keeps local model context within a range the GPU can handle well.
- Compaction policy: Compacts when context reaches a fixed proportion of the model-declared context window.
- Multi-dimensional checks:
hardCeilingTokens: hard ceiling check.maxTtftMs: measures Time to First Token; triggered if the configured value is exceeded.maxColdPrefillMs: predicts cold prefill time, i.e. the time required to prefill the current context if the server cache is lost.
- Action execution: supports
warn(print warning only) orcompact(perform compaction). - Monitoring command: provides the
/context-budgetcommand to view current values, historical measurement samples, and predicted cold prefill time.
Installation and Enablement¶
First, use the officially provided command to install the plugin. The plugin is loaded through DSH’s package manager.
dsh plugin --profile web add dsh-context-budget
After installation, edit the configuration file. The configuration file path is typically ~/.dsh/profiles/web/cordis.patch.yml. Add a context-budget configuration block to this file.
- id: context-budget
config:
providers:
llamacpp:
hardCeilingTokens: 110000 # 可选,硬性上下文 token 上限
maxTtftMs: 180000 # 可选,首字延迟上限,单位毫秒
maxColdPrefillMs: 600000 # 可选,冷预填充时间上限,单位毫秒
retainTokens: 24000 # 压缩时保留的 token 数量(默认 16000)
action: warn # 动作:warn(默认)或 compact
After configuration, restart dsh web and open a new session. You can verify that the configuration took effect using dsh --profile web --dump-config.
Typical Usage¶
After the plugin is configured, it runs in the background. When a configured check is triggered, it prints logs or performs compaction according to the action setting.
To view the current context state, historical samples, and predicted cold prefill time, enter the command directly in the session:
/context-budget
The example output shows the current context token count, trigger status of each check, recent prefill rate samples including cold/hot state, predicted cold prefill time, and estimated time required for compaction.
Use Cases and Notes¶
This plugin is suitable for running large-parameter models locally to avoid severe performance degradation caused by overly long context.
- Ecosystem coexistence: The plugin coexists with DSH’s built-in
compaction-basic. If the agent preset for the session does not have a compaction engine configured, thecompactaction degrades towarn. - Measurement scope: The compaction summarization request itself is not counted in measurements. Observed slow responses are forgotten after the context is rewritten, so a single compaction does not trigger a chain reaction.
- Log output: Logs are printed only when checks are triggered. If the session always stays within the limits, the terminal will show no output.
- Source note: This plugin is community-maintained and is not an official DeepSeek product. See the links below for source code and repository information.
Summary¶
dsh-context-budget provides a context management approach based on measured data. By combining soft and hard check mechanisms, it effectively addresses performance degradation of local models under long context.