DeepSeek Harness (DSH) follows the “everything is a plugin” architectural philosophy. For agent development based on pure text models, image recognition capability is often a weakness. The dsh-quicksight plugin uses a two-tier strategy of “local OCR + vision model” to add image recognition capability to pure text models.
The plugin is maintained by Isanti2016 and licensed under the MIT License. It registers a tool named quicksight_ocr, which can automatically decide based on image content whether to invoke a local OCR engine or fall back to a vision model for structured analysis.
How It Works¶
The plugin’s core logic is “fast first, then accurate”: first try to extract text quickly using local OCR; if the result does not meet the standard (e.g., insufficient characters or garbled output), then invoke a vision model for deeper analysis.
图片输入
│
▼
quicksight_ocr 工具
├─ Tier-1: 本地 RapidOCR
│ └─ 文本 ≥ minChars(默认20) → 直接返回文本
│ (耗时约 2-3 秒,完全离线,零 API 成本)
└─ Tier-2: modlens 视觉模型
└─ OCR 失败或需要结构化证据 → 调用视觉引擎
(耗时约 20 秒,返回 summary/OCR/版面/语义)
Installation & Dependencies¶
1. Install the Plugin¶
Use the official installation command to add the plugin. After installation, restart the dsh web service for the tool to take effect.
npx -y @deepseek-ai/dsh plugin --profile web add dsh-quicksight
2. Prepare Tier-1 Local OCR (Recommended)¶
Tier-1 depends on a Python environment and RapidOCR. It is recommended to create a separate virtual environment.
# 创建虚拟环境
python -m venv ~/.dsh/tools/ocr-venv
# 安装 RapidOCR
~/.dsh/tools/ocr-venv/Scripts/pip install rapidocr_onnxruntime # Windows
~/.dsh/tools/ocr-venv/bin/pip install rapidocr_onnxruntime # macOS/Linux
3. Prepare Tier-2 Vision Engine (Optional)¶
Tier-2 reuses the modlens ecosystem. You need to install the modlens plugin and configure at least one vision engine (e.g., Nvidia NIM).
npx -y @liustack/modlens config set openai.baseUrl <https://.../v1>
npx -y @liustack/modlens config set openai.apiKey <你的 key>
npx -y @liustack/modlens config set openai.model <视觉模型>
npx -y @liustack/modlens config set provider openai
Configuration¶
The plugin can be configured via environment variables or cordis.patch.yml. The main configuration items are as follows:
| Key | Default | Description |
|---|---|---|
toolName |
quicksight_ocr |
Registered tool name |
minChars |
20 |
Tier-1 success threshold for extracted text (character count) |
ocrPython |
~/.dsh/tools/ocr-venv/{Scripts\|bin}/python |
Path to the Python interpreter with rapidocr installed |
modlensEnabled |
true |
Whether to enable the Tier-2 fallback mechanism |
timeoutMs |
120000 |
Timeout for a single recognition request (milliseconds) |
Runtime & Security¶
- Port listening: The plugin does not listen on any port. It is a pure tool plugin and does not expose HTTP/WS services.
- Privacy policy:
- Tier-1: Images are processed entirely locally, are not uploaded to any server, and incur zero API cost.
- Tier-2: Images are sent to the user-configured vision engine (e.g., Nvidia NIM, Gemini, etc.).
- Credential management: API Keys are stored only locally in
~/.modlens/config.json, are not included in the plugin code, and are not written to session logs.
Known Limitations¶
- Free engine limitations: If using a free vision engine (e.g., one with limited free quota), you may encounter rate limits (returning 429/404). Tier-2 failures will be reported as a degraded error.
- Capability boundaries: Tier-1 only extracts text and does not support analysis of colors, layouts, charts, or faces; such analysis depends on the Tier-2 vision model.
- Format support: Supports common image formats (PNG/JPG/BMP/TIFF/WebP), but does not support PDF documents.