Introduction¶
The core design philosophy of DeepSeek Harness (DSH) is “everything is a plugin.” When building long-conversation agents, cloud-side large models need to reread the entire conversation history every turn. As the history grows longer, the density of useful information is diluted. Context Assembler DSH addresses this fundamental problem: it uses idle local compute cycles to reassemble the context for each turn. By organizing context into topic blocks, it keeps details in relevant areas, downgrades the rest into structured summaries, and injects background information when topics shift, thereby maximizing the mutual information density per token.
Plugin Positioning¶
Context Assembler DSH is a pure computation plugin for DeepSeek Harness (dsh), maintained by developer i1j. It does not depend on the host environment, but it needs to run with local small models.
Core Features¶
-
Topic Block Context Assembly
Real-time splitting of the session into topic blocks, rebuilding the summary history from the perspective of the current block. Within a topic block, the summary prefix remains stable, which helps with cache hits; the tail retains a raw-data protection zone (tailN), ensuring current-turn information is not lost. -
Water Pressure Topic Splitting
Based on Hermes’applyWaterPressurealgorithm. As context characters accumulate, the Jaccard similarity threshold required for splitting is dynamically lowered. Splitting is forced at peak pressure to prevent a single topic from becoming too long and causing cache invalidation. -
Topic Grading and Freezing
Freeze ACT/REL/FAR levels when a topic switch occurs. New turns remain at the ACT level (deterministic, no LLM call required) until the next switch. -
Tool Turn Compression
Structurally summarizetoolCallandtoolResult(deterministic, no LLM required). When the token saving threshold is reached, rewrite tool results line by line. -
Reality Recall Injection
Use a local 4B-class embedding model for retrieval and inject relevant backgrounds when a topic block begins. This feature follows a fail-open design: if the database is missing, only the feature is disabled without raising an error. -
Thought (OODA) Assembly
Perform multi-transaction assembly for thought streams and tool streams, and extract L1-level factual appendices from local 4B offline cards. -
Handoff Planning
Pressure-triggered session handoff, calculating branch summaries, edge strength, perspective, and routing strategy. -
ca-dbCommon Library
Provide external database support.
Environment Dependencies¶
- Zero host dependency: The plugin core itself does not depend on the host environment.
- Local small model: Depends on local small models for inference. The model must expose an Ollama-compatible interface; a parameter scale of 4B is recommended.
- Hardware recommendation: A 10GB VRAM GPU is recommended to achieve good performance (it can smoothly run a quantized 4B model). Without a GPU, CPU inference can be used, but it is slower.
Notes¶
- Must be used together with DeepSeek Harness (dsh).
- The code is open source and released under the MIT License.
- It is recommended to review the source code and license before use.
Conclusion¶
This plugin replaces high-frequency cloud bandwidth operations with local computation, aiming to solve context management challenges in long-conversation scenarios. For more details, see the GitHub repository or the community catalog.