Introduction¶
In long conversation scenarios, the context window grows linearly, eventually causing the model to overflow. Traditional Threshold Summarization approaches usually create “memory interruptions” at compression time, or misclassify the latest instructions as redundant content.
Mosaic Memory Compress provides a stateless conversation compression solution. It mimics the human memory forgetting curve, keeping conversation continuity while controlling the number of messages within a fixed range. The plugin includes a DeepSeek Harness (DSH) adapter and can be used directly within the DSH workflow.
What Is It?¶
This is a general-purpose, pluggable, stateless conversation compression algorithm that supports any LLM agent framework. It includes a ready-to-use DeepSeek Harness (DSH) adapter module. With this plugin, LLM conversations can remain permanently bounded without complex session management and without worrying about context overflow.
Core Features¶
- Stateless and repeatable: No dependency on session state. Each call returns deterministic results, and the results can be used directly as input for the next round.
- Zero cost below the threshold: If the compression threshold has not been reached, the function returns immediately and does not consume extra resources.
- Anti-jitter: Compression is triggered only at configured window boundaries, avoiding performance jitter caused by high-frequency triggering.
- LLM-agnostic: Supports any LLM provider. By customizing the
callLLMfunction, it can connect to OpenAI, Anthropic, or local models. - DeepSeek Harness (DSH) adapter: Provides a dedicated module for seamless integration with DeepSeek Harness.
- Tool-call safe: In messages containing tool calls, it does not break turn-counting logic.
- Graceful degradation: Even if an LLM call fails, it does not block the conversation flow.
Installation and Enablement¶
Install the core package via npm:
npm install mosaic-memory-compress
The plugin is processed using cordis.patch.yml during packaging.
Typical Usage¶
Import the function and configure the compression parameters. MosaicMemoryConfig includes threshold parameters that control the compression cadence and an LLM callback function for handling heavyweight compression.
import { mosaicMemoryCompress, type MosaicMemoryConfig } from 'mosaic-memory-compress';
const config: MosaicMemoryConfig = {
lightStart: 10, // 保持最近 10 轮原始消息
lightWindow: 30, // 每 30 轮进行一次轻量级压缩
heavyStart: 40, // 超过此轮数进入重载区
heavyWindow: 30, // 重载压缩的节奏
callLLM: async (systemPrompt, userInput) => {
// 在这里接入你的 LLM 提供商(如 OpenAI、DeepSeek 等)
// const res = await openai.chat.completions.create({...});
// return res.choices[0].message.content ?? '';
return ''; // 占位
},
};
// 在对话的每一轮调用此函数
const compressed = await mosaicMemoryCompress(messages, config);
Applicable Scenarios and Considerations¶
- Applicable scenarios: Scenarios requiring processing of unlimited-length conversations, or developers who want to avoid the “memory interruption” problem in traditional summarization approaches.
- Important notes:
- No session state: This plugin is completely stateless and is not responsible for saving context. Persistent storage is the host’s responsibility (for example, storing the compressed array in a file or database).
- License check: Check the source code and license (MIT) before use.
- DSH permissions: The plugin runs with the current DSH process permissions. Ensure that you can review the source code.
Summary¶
Mosaic Memory Compress solves the pain point of long-conversation context overflow through a design of “statelessness + threshold control + model agnosticism”. It provides a DSH adapter, allowing conversations to remain permanently bounded while preserving “living memory”.