Introduction¶
DeepSeek Harness (DSH) uses a plugin-based architecture that allows users to extend functionality as needed. For developers deploying local Qwen3.8-series models (27B, Flash-Next), managing Thinking Budgets across different inference backends and handling context window Compaction is a common pain point. The dsh-qwen38-local-qol plugin aims to provide a unified configuration entry point and optimization strategy for this local workflow.
Plugin Overview¶
This is a Quality of Life (QoL) plugin for DeepSeek Harness, maintained by the user Yunado and licensed under the MIT License. It mainly addresses the configuration fragmentation problem when running Qwen3.8 on local inference engines such as llama.cpp, NInfer, TabbyAPI, and oMLX.
Core Features¶
- Per-request Thinking Budget: Supports four dialects—llama.cpp, NInfer, TabbyAPI, and oMLX—allowing the reasoning budget to be adjusted dynamically in each request.
- Compaction backend optimization: When compaction is triggered, it trims prefill content (preserves recent reasoning, downgrades images to text placeholders, and truncates tool results), preventing checkpoint truncation caused by output limits.
- Visual support in tool results: Image blocks nested in tool results (such as
read_imageoutput) are transmitted rather than discarded. - Visual settings panel: Provides a Settings Tab, supporting real-time configuration of all row parameters.
Installation and Enablement¶
Install the plugin from the command line:
dsh plugin --profile web add github:Yunado/dsh-qwen38-local-qol
After installation, the dsh web service must be restarted. When the service starts, the plugin generates a qwen38 user preset from the standard preset combination. If no default agent preset is configured, this preset becomes the default. New sessions automatically use this preset, while existing sessions retain the preset they were created with.
To pin a specific version, add a tag to the specification (such as #v0.2.0 or any previous release version).
Configuration Details¶
The plugin adds a Settings Tab named Qwen3.8 Local in DSH settings. This tab lists all configuration items and supports real-time modification and hot reloading.
Status Indicator¶
The status dot color in the settings panel indicates the default status of the current qwen38 preset:
* Green: qwen38 is the default preset.
* Amber: The default preset is another preset other than qwen38.
* Gray: The qwen38 preset does not exist.
Server Configuration¶
This Tab supports configuring four different local inference servers:
- llama.cpp (
llamacpp): The standardllama-server, which supports sending Thinking Budgets dynamically through request parameters. - NInfer: A high-performance TensorRT-LLM engine, where the Thinking Budget is typically configured at server startup (
--default-thinking-budget). - TabbyAPI: An ExLlamaV3-based backend that natively supports per-request budgets.
- oMLX: An Apple Silicon MLX inference server that natively supports
thinking_budgetandchat_template_kwargs.enable_thinking.
The main configuration items include:
* Dialect (dialect): Selects one of the engines above (default llamacpp).
* Base URL: The local service address (default ports are: llama.cpp 8080, NInfer 8082, TabbyAPI 8083, oMLX 8000).
* Model: The model ID or alias recognized by the server.
* Display Name: The label shown in the UI model selector.
* API Key: An optional authentication token.
Context Window and Output Limits¶
- Context Window: The total Token capacity supported by the model, used to determine when to trigger compaction.
- Max Tokens: The maximum Token count for a single response, default to about 20% of the context window (approximately 52,428 Tokens), reserving space for compaction triggering.
Thinking Budget¶
The plugin supports setting different Thinking Budget levels for llama.cpp, TabbyAPI, and oMLX:
* Low: 4,096 Tokens
* Medium: 8,192 Tokens
* Extra High: 16,384 Tokens
For NInfer, configure it through defaultThinkingBudget. There is also a defaultEffort parameter for setting the default thinking intensity.
Compaction and Trimming¶
To prevent context overflow, the plugin provides a set of compaction strategies (configured mainly at the Env variable level, but also configurable in the Tab):
* Summarize Images: The strategy for handling historical images. strip (recommended) replaces older images with text placeholders to save space; keep retains the images.
* Keep Recent Reasoning: Preserves the several most recent assistant replies.
* Tool Chars: Limits the maximum number of characters for tool results.
Notes¶
- Dependency requirement: The plugin requires a Node.js environment version >= 22.14.
- No core patches: This plugin does not include core patches or a pi-ai patchfile.
- Persistence: All changes made in the settings panel take effect in real time and are persisted to the
settings.yamlfile.