Introduction¶
DeepSeek Harness (DSH) uses context compression to save tokens in long conversations. However, the locally deployed qwen3.8-27b model (especially when running on the NInfer engine) defaults to an xhigh reasoning level. As a result, the model exhausts its output token budget before generating a summary, causing the compression checkpoint to be truncated and contextual information to be lost.
What Is This¶
This is a DSH plugin maintained by zhubaohi, used to fix compression failure issues for the qwen3.8-27b gateway served by Neroued’s NInfer engine. The plugin adjusts the reasoning strategy to ensure that compression and title generation complete successfully.
Core Features¶
- Disable thinking to save budget: Disable thinking only for context compression and session title generation calls, preventing the
xhighreasoning level from exhausting the output budget. - Apply recommended sampling parameters: Apply model-recommended sampling parameters (such as
temperature: 0.7,top_p: 0.8, etc.) in non-thinking mode. - Raise output limit: Raise
max_tokensfor compression requests to a configurable floor (default 16384), preventing over-compression of context. - Fix title generation: Fix title generation truncation by controlling the reasoning intensity of title requests.
- Scope limitation: Applies only to NInfer engine gateways.
- Strict model matching: Requires a matching rule exactly matching the model ID.
Install and Enable¶
Install the plugin via CLI. After installation, restart dsh web or refresh the GUI page for changes to take effect.
dsh plugin --profile web add dsh-qwen38-ninfer-compaction-fix
Configuration and Usage¶
The plugin requires the target model ID, sampling parameters, and output limit to be specified in configuration files.
1. Configure settings.yaml¶
Add the configuration in $DSH_HOME/settings.yaml. This configuration supports hot loading (no restart required).
qwen38-compaction-fix:
effort: off # 推理强度设为 off
models: [qwen3.8-27b] # 精确匹配的模型 ID
sampling: # 写入请求体的采样参数
temperature: 0.7
top_p: 0.8
top_k: 20
min_p: 0.0
presence_penalty: 1.5
repetition_penalty: 1.0
maxTokensFloor: 16384 # 输出 token 下限
titleReasoning: none # 标题生成推理强度
2. Configure cordis.patch.yml¶
It can also be configured in cordis.patch.yml. This takes effect the next time the GUI loads.
- id: qwen38-compaction-fix
config:
models: [qwen3.8-27b]
Notes¶
- Engine limitation: The plugin is only compatible with NInfer engine gateways. If started with other engines such as llama.cpp or vLLM, although the approach is the same, parameters must be adapted manually.
- Restart required: After installation, you must restart the
dsh webprocess or refresh the UI. - Model ID matching: The plugin performs exact matching on the
modelfield in requests. If the configured ID does not match (for example, case differences), the plugin silently lets the request pass through without modifying it. If you encounter compression truncation, first check whether the model ID configuration is correct.