Preface¶
DeepSeek Harness (DSH) is a plugin-based ecosystem whose core lies in context window management. Cloud models (such as DeepSeek V4) have a million-token window, and session recovery via resume plus history compaction via compaction are sufficient to maintain memory. However, once local models (Ollama, etc.) are connected, the window is typically only 4K-32K, making it impossible to fit the entire history into the prompt.
At this point, a structured persistence layer is needed to extract and store facts from conversations that are worth retaining, and inject them on demand in subsequent conversations. dsh-memory-bridge was built for exactly this scenario: it provides a retrievable, auditable, and governable long-term memory system when the context window is constrained.
What is It¶
This is a memory-tree bridging plugin. It automatically distills conversations into a searchable, visualizable memory tree, and injects relevant context into subsequent sessions on demand. The project is positioned as “memory infrastructure” rather than a perfect agent. By combining a plaintext Markdown source of truth with a SQLite index, it achieves transparent, controllable data and fast retrieval.
Core Capabilities¶
The plugin provides the following core features:
- Cross-session persistent memory: Conversations are automatically distillled into memory cards, which are retrieved and injected on demand in subsequent sessions, supplementing DSH’s native session history with an additional persistent fact layer.
- Memory retrieval tools: Agents can actively read and write memory using the
memory_search,memory_add_run, andmemory_reviewtools. - Memory visualization: The settings page includes 7 visualization tabs, covering event graph, knowledge graph, timeline, pending review, profile, audit, and overview.
- Experience and profile distillation: Automatically identifies “remember the lesson / learn from the pitfall” patterns and distills them into permanent experiences; aggregated preference signals can trigger profile distillation.
- Governance loop: Prevents unbounded memory expansion through forgetting curves (30-day idle expiration), utilization-based contraction, and audit feedback.
- Storage and retrieval: Uses a plaintext Markdown source of truth + SQLite index; retrieval uses zero-LLM algorithms (BM25 + RRF).
- Dual-channel writing: Supports LLM extraction (breadth) and zero-LLM rules (speed), ensuring critical information can be persisted immediately.
Installation and Dependencies¶
The plugin is distributed as an npm package, maintained by gangwolf2312-creator. Before installing, make sure jieba is installed in the local environment.
Dependencies:
- @deepseek-ai/dsh-tools: ^0.1.0-rc.7
- @deepseek-ai/cordis: ^4.0.1 (peer dependency)
Configuration:
The plugin relies on the environment variable apiKeyEnv. It is recommended to configure a backend LLM API key to enable extraction features.
Typical Usage¶
- Agent invocation: Agents can call
memory_searchdirectly to query historical memory, or usememory_add_runto add new memories. - Injection strategy: The plugin injects memories according to intent tiers.
- Critical intent: up to 3 memories.
- General intent: up to 1 memory.
- Greeting / small talk: no memory injection.
- Profile distillation: Manually trigger in the UI’s “Profile” tab, or call
distillvia RPC, to aggregate the event tree into a user profile summary. - Review process: Uncertain content from zero-LLM extraction enters the
pendingreview queue and must be confirmed by a human before being finalized.
Design and Implementation¶
- Storage layer: Each memory card is a Markdown file (including front matter metadata and body), and directories are organized by type (e.g.,
events/cards,profiles). Write operations are idempotent, and reconciliation is possible after a crash and restart. - Retrieval mechanism: The read side has zero LLM dependencies. It uses jieba tokenization to extract word tags and computes relevance through multi-source RRF fusion. Data-driven relevance thresholds (such as core-term overlap thresholds) prevent injection of irrelevant memories.
- Memory-knowledge separation: The system distinguishes between
memory-tree(experiences about “you”) andmemory-wiki(normative knowledge about “the world”), avoiding retrieval pollution. - Architecture isolation: Uses a Host JS + Python Sidecar pattern. The Sidecar only listens on 127.0.0.1, and a crash only affects the memory module without bringing down the main Harness process.
Caveats¶
- LLM dependency: Extraction quality completely depends on the output capability of the selected LLM.
- jieba installation: Ensure
jiebais installed locally before use. - Retrieval limitation: Retrieval is based on relevance-ranked recall rather than semantic association, with limited expansion capability.
- Fault isolation: A Sidecar crash only makes memory read/write unavailable; it does not affect Harness session flow.
- License: MIT License, open-source agreement. Please comply with it when using.
Summary¶
For developers with constrained context windows or users of local models, dsh-memory-bridge provides a reliable memory governance solution. Through zero-LLM retrieval, dual-channel writing, and plaintext storage, it addresses the pain points of agent “forgetfulness” and “context overflow” without sacrificing controllability.
Project URL: https://github.com/gangwolf2312-creator/dsh-memory-bridge