The DSH plugin ecosystem is typically used to extend the capabilities of DeepSeek Harness. In voice transcription scenarios, transcripts often contain many filler words, repeated content, and politeness phrases, which consume a large number of tokens and affect context understanding. dsh-voice-prompt-compressor is designed to address this pain point by providing a deterministic, local text compression solution.
Positioning¶
The plugin is maintained by developer yuzh1090 under the MIT License. It is a DeepSeek Harness (DSH) plugin whose core function is to compress lengthy voice transcripts into token-efficient prompts. The entire process runs completely locally, makes no network requests, and consumes no LLM tokens.
Core Features¶
- Local deterministic compression: Compresses text through normalization, filler removal, politeness removal, and deduplication. The entire process is offline, does not call an LLM, and consumes no tokens.
- Bilingual lexicons: Built-in Chinese and English word lists for filler words, hesitation markers, and politeness phrases; supports automatic language detection based on the CJK character ratio.
- Tool and skill support: Registers the
compress_voice_texttool and thevoice-prompt-compressorskill. The Agent can call the tool directly or trigger it automatically through the skill. - Configurable behavior: Supports adjusting the compression mode (
mode) and whether to preserve politeness phrases (keepPoliteness) through configuration options.
Installation¶
Use the official installation command:
dsh plugin --profile web add dsh-voice-prompt-compressor
Note: The dist/ directory in the source repository is not committed. Before installing, clone the repository and build it first:
npm install
npm run build
After installation, refresh the Web page. The plugin takes effect in new sessions.
Usage¶
Skill Mode (Recommended)¶
When a user pastes a lengthy voice transcript into the chat, the Agent automatically loads the voice-prompt-compressor skill, calls compress_voice_text, and organizes the compressed text into a structured prompt with four sections (Context / Goal / Constraints / Deliverables).
Tool Mode¶
The Agent can directly call the compress_voice_text tool with the following parameters:
text(string, required): The voice transcript text to be compressed.language(string, optional): The lexicon language. Defaults toauto(detected by CJK ratio).mode(string, optional): The compression strength. Defaults to the configuration value. Accepted values:light,balanced,aggressive.keepPoliteness(boolean, optional): Whether to preserve politeness phrases. Defaults to the configuration value.
Example Return Value:
{
"compressed": "…",
"originalLength": 512,
"compressedLength": 210,
"estimatedTokensSaved": 76,
"ratio": 0.59,
"removedCategories": {
"fillers": 18,
"repeats": 3,
"politeness": 2
}
}
Technical Details¶
The plugin uses a purely mechanical, deterministic processing pipeline:
- Normalization: Converts full-width characters to half-width and normalizes whitespace.
- Filler removal: Removes filler words and discourse markers (e.g., “uh”, “um”, “you know”, “I mean”, “like”, etc.). In
balancedmode, ambiguous pronouns are removed only when near punctuation or whitespace; inaggressivemode, more pronouns and hesitation markers (e.g., “to be honest”) are removed. - Politeness removal: Removes polite filler phrases (e.g., “please”, “thanks”, etc.), unless preservation is configured.
- Deduplication: Merges adjacent repeated content (e.g., “no no” → “no”).
The compression process performs only text processing and does not rewrite the original meaning. Technical requirements, constraints, and business rules are preserved verbatim.
Development¶
The plugin requires a Node.js environment (>= 18). Main dependencies include @deepseek-ai/schemastery and yaml. Development involves building the TypeScript project: use npm install to install dependencies, use npm run build to produce the build artifacts, and use npm test to run tests.