Introduction

When DeepSeek Harness (DSH) processes long text, input token consumption can increase rapidly, leading to higher costs or context window truncation. This plugin compresses input tokens by rendering long text as images and sending them, leveraging the per-image token cap of 384 tokens used by vision models. It is designed specifically for DSH, requires no custom RPC, and integrates into DSH’s existing architecture.

Core Features

  1. Text-to-Image Compression: Renders long text into 800×800 images and uses the 384-token per-image cap to compress LLM input. The input seen by the model drops from tens of thousands of tokens to hundreds or thousands, while the content remains word-for-word unchanged.
  2. Automatic Scaling and Compression Ratio: Supports automatic image scaling. For Chinese text at 18px, the compression ratio is about 2~2.5×; at 13px it is higher. For English text at 18px or larger, compression provides almost no benefit.
  3. Smart Model Detection: Automatically detects whether the model is a vision model. Non-vision models (those not containing “vision”) automatically skip conversion to avoid errors.
  4. Architecture Integration: The Host listens to the agent/pre-step event to replace text blocks, while the Client injects the “Image” toggle and font size dropdown. The toggle and font size are managed through settings; rendering and replacement use host events and the attachment service, with zero custom RPC.
  5. Threshold and Page Count Control: Supports configuring a threshold (default 600 characters) and a maximum page count (default 10 pages, range 1–20). If limits are exceeded, it automatically falls back to plain text.

Installation and Enablement

Install the plugin using the DSH CLI:

dsh plugin --profile web add dsh-text2img-compress@0.1.3

After installation, restart DSH and refresh the page. An “Image” button will appear on the right side of the input box.

Usage

  1. Enable the Feature: Click the “Image” button on the right side of the input box. The state is persisted and is not lost after restart.
  2. Send Long Text: Send text of ≥ 600 characters (default). The chat interface displays a prompt and automatically renders the text as images.
  3. Adjust Parameters: In the settings panel (DSH settings panel → “Text-to-Image Compression” page), adjust the font size (12–22px, in 1px increments) and the maximum page count (1–20) to control the compression ratio.
  4. Model Reading: The model reads the image content directly, enabling high-fidelity reading or summarization.

Use Cases and Limitations

Use Cases:
- Understanding, summarizing, and retrieving content from long documents, papers, manuals, or release notes.
- Budget compression for long context windows (for example, reducing a 100KB document from 30,000+ tokens to a few thousand).
- Scenarios requiring approximate reconstruction of content (dates, numbers, paragraph-level meaning).

Limitations and Notes:
- Vision Models Only: Applies only to vision models (such as deepseek-v4-flash-vision-exp). Pure text models are skipped automatically.
- English Scenarios: Compression for English text at 18px or larger provides almost no benefit. Use 13px or plain text instead.
- Precision Requirements: Not suitable for scenarios requiring word-for-word precision, such as code, configuration, SQL, or JSON.
- OCR Accuracy: General-purpose vision models are not OCR-specific models. Use 22px when high accuracy is required.

Conclusion

By leveraging the per-image token cap mechanism of vision models, this plugin effectively reduces input token consumption for long text processing. Installation and configuration are simple, making it suitable for scenarios that require efficient long-text handling in DSH.