DeepSeek Harness (DSH) follows the “everything is a plugin” architectural philosophy. For agent development based on pure text models, image recognition capability is often a weakness. The dsh-quicksight plugin uses a two-tier strategy of “local OCR + vision model” to add image recognition capability to pure text models.

The plugin is maintained by Isanti2016 and licensed under the MIT License. It registers a tool named quicksight_ocr, which can automatically decide based on image content whether to invoke a local OCR engine or fall back to a vision model for structured analysis.

How It Works

The plugin’s core logic is “fast first, then accurate”: first try to extract text quickly using local OCR; if the result does not meet the standard (e.g., insufficient characters or garbled output), then invoke a vision model for deeper analysis.

图片输入
    │
    ▼
quicksight_ocr 工具
    ├─ Tier-1: 本地 RapidOCR
    │   └─ 文本 ≥ minChars(默认20) → 直接返回文本
    │      (耗时约 2-3 秒,完全离线,零 API 成本)
    └─ Tier-2: modlens 视觉模型
       └─ OCR 失败或需要结构化证据 → 调用视觉引擎
          (耗时约 20 秒,返回 summary/OCR/版面/语义)

Installation & Dependencies

1. Install the Plugin

Use the official installation command to add the plugin. After installation, restart the dsh web service for the tool to take effect.

npx -y @deepseek-ai/dsh plugin --profile web add dsh-quicksight

Tier-1 depends on a Python environment and RapidOCR. It is recommended to create a separate virtual environment.

# 创建虚拟环境
python -m venv ~/.dsh/tools/ocr-venv

# 安装 RapidOCR
~/.dsh/tools/ocr-venv/Scripts/pip install rapidocr_onnxruntime   # Windows
~/.dsh/tools/ocr-venv/bin/pip install rapidocr_onnxruntime       # macOS/Linux

3. Prepare Tier-2 Vision Engine (Optional)

Tier-2 reuses the modlens ecosystem. You need to install the modlens plugin and configure at least one vision engine (e.g., Nvidia NIM).

npx -y @liustack/modlens config set openai.baseUrl <https://.../v1>
npx -y @liustack/modlens config set openai.apiKey <你的 key>
npx -y @liustack/modlens config set openai.model <视觉模型>
npx -y @liustack/modlens config set provider openai

Configuration

The plugin can be configured via environment variables or cordis.patch.yml. The main configuration items are as follows:

Key Default Description
toolName quicksight_ocr Registered tool name
minChars 20 Tier-1 success threshold for extracted text (character count)
ocrPython ~/.dsh/tools/ocr-venv/{Scripts\|bin}/python Path to the Python interpreter with rapidocr installed
modlensEnabled true Whether to enable the Tier-2 fallback mechanism
timeoutMs 120000 Timeout for a single recognition request (milliseconds)

Runtime & Security

  • Port listening: The plugin does not listen on any port. It is a pure tool plugin and does not expose HTTP/WS services.
  • Privacy policy:
    • Tier-1: Images are processed entirely locally, are not uploaded to any server, and incur zero API cost.
    • Tier-2: Images are sent to the user-configured vision engine (e.g., Nvidia NIM, Gemini, etc.).
  • Credential management: API Keys are stored only locally in ~/.modlens/config.json, are not included in the plugin code, and are not written to session logs.

Known Limitations

  • Free engine limitations: If using a free vision engine (e.g., one with limited free quota), you may encounter rate limits (returning 429/404). Tier-2 failures will be reported as a degraded error.
  • Capability boundaries: Tier-1 only extracts text and does not support analysis of colors, layouts, charts, or faces; such analysis depends on the Tier-2 vision model.
  • Format support: Supports common image formats (PNG/JPG/BMP/TIFF/WebP), but does not support PDF documents.