Introduction¶
The DeepSeek Harness (DSH) plugin ecosystem emphasizes modularity and flexibility. When building long-conversation agents, context window limits are a common bottleneck, and simple truncation or summarization often makes it difficult to balance information density with computational cost. dsh-infinite-context addresses this problem by providing a complete multi-layer memory management solution.
Plugin Overview¶
This is a DeepSeek Harness (DSH) plugin designed to provide an “infinite context” experience for long conversations through multi-layer memory management, semantic retrieval, and model context awareness.
- Maintainer: chocobo77
- License: MIT
- Core Value: Solves context overflow in long conversations while maintaining dialogue coherence and retention of key information.
Core Features¶
The plugin maintains long conversations through the following mechanisms:
- Multi-layer memory pyramid: Divides memory into three layers — short (recent raw content), mid (LLM summaries), and long (consolidated summaries) — to manage information of different timeliness and importance.
- Progressive compression and dynamic thresholds: Dynamically adjusts compression trigger levels based on token usage. When the context exceeds the threshold, older messages are automatically summarized, while recent conversations are preserved as-is.
- Deep thinking intervention: Wraps the LLM stream and injects an overflow signal when inputs or outputs approach the actual model context window, triggering compression and retry, so intervention can still occur during the model’s thinking process.
- Semantic retrieval and three-level deduplication: Uses embedding vectors for semantic retrieval and applies exact match, normalized fuzzy match, and semantic cosine similarity (threshold 0.92) for deduplication to prevent duplicate information from being stored.
- Structured memory and persistence: Memory is categorized by user/feedback/project/reference and stored in SQLite, supporting restoration after restart.
- Model context awareness: Automatically reads the actual model context window resolved by DSH and supports local model probing, ensuring compression is triggered within the real window rather than using declared values.
Installation and Deployment¶
The installation command is:
dsh plugin add
According to the documentation, this installation method has a temporary-path trap. dsh plugin --profile <name> add <tgz> stages the tarball in a temporary directory, which may be cleaned up by the system and cause subsequent installation failures. The recommended approach is to install using a persistent directory, or follow the detailed deployment instructions in the README (for example, package it with npm pack and place it in ~/.dsh/packages/).
Configuration¶
The configuration is mainly divided into memory-context and memory-compaction sections.
memory-context Configuration¶
| Key | Default | Description |
|---|---|---|
storePath |
dsh-infinite-context.db |
SQLite database storage path. Set to :memory: to disable persistence. |
contextWindow |
94000 |
Model context window (fallback). The plugin prioritizes the actual window resolved by DSH. |
headroomRatio |
0.25 |
Ratio reserved for system, tools, inputs, and outputs. |
modelProbe.enabled |
false |
When enabled, actively probes the actual context window of local servers (e.g., llama-server/Ollama). |
budget.short/mid/long/retrieved |
10000/20000/5000/15000 |
Layered token budget. |
memory-compaction Configuration¶
| Key | Default | Description |
|---|---|---|
thresholdRatio |
0.7 |
Compaction trigger ratio. Basic compaction is triggered when context usage exceeds 70% of the window. |
compaction_dynamic_threshold |
true |
When the real window is smaller than the declared window, derive the threshold from the real window and force history compaction. |
thinking_guard_enabled |
true |
Enable deep thinking intervention and monitor overflow in the generation stream. |
compress_target_ratio |
0.6 |
Target compaction level; only processes the overflow portion. |
retainRatio |
0.3 |
Retain the most recent 30% of window content during compaction. |
Manual Tools¶
The plugin provides 10 manual tools for managing memory:
memory_search(query?, k?): Semantic search persisted memories.memory_status: Reports layered counts, budgets, embedder, forgetting policy, and model context.memory_index(limit?): Generates a structured index in MEMORY.md style.memory_maintain: Runs a read-only audit for duplicates, conflicts, and outdated information.memory_model_probe(forceProbe?, model?): Reports the model context source and can force probing.memory_forget: Runs a forgetting scan (memories belowforgetting.minScore).memory_consolidate: Forces pyramid consolidation.memory_reset: Clears all memories.memory_force_compress(sessionId?): Forces compression for a specified session.memory_ingest(text, source): Manually ingests a text entry.
Use Cases and Notes¶
- Dependency environment: Depends on the DeepSeek Harness plugin family and requires Node.js >= 22.19.0.
- Permissions: The plugin runs with the permissions of the current DSH process. Review the source code and license before installation.
- Test coverage: The project includes 133 unit tests.
Conclusion¶
dsh-infinite-context provides DSH plugins with practical long-conversation management capabilities through layered memory, dynamic compression, and model awareness. For developers who need to handle ultra-long contexts or complex multi-turn conversations, this plugin can effectively alleviate context overflow issues.
- Catalog page: https://www.skillhub.cn/plugins/chocobo77/dsh-infinite-context
- GitHub repository: https://github.com/chocobo77/dsh-infinite-context