DSH adopts the “everything is a plugin” philosophy. When processing long documents, tables, or complex layouts, we need a clear visual processing pipeline and auditability. duhai-vision is a general-purpose vision model adapter that supports Codex and DSH, aiming to provide replaceable visual processing capabilities without affecting visual quality.

Core Positioning

  • Function: General-purpose vision model adapter.
  • Supported targets: Codex and DeepSeek Harness.
  • Underlying model: PaddleOCR-VL is preferred by default; Qwen is used as a fallback or when explicitly selected.
  • Maintainer: hamliy-feng.
  • License: MIT.

Installation and Configuration

Before use, you need to install runtime dependencies and configure access credentials.

1. Install Dependencies

Install the PaddleOCR-VL runtime dependencies:

python -m pip install "paddleocr>=3.4,<4"

2. Configure Paddle Access Token

Configure the Paddle Access Token:

$env:PADDLEOCR_ACCESS_TOKEN = "<你的 Access Token>"
[Environment]::SetEnvironmentVariable(
  "PADDLEOCR_ACCESS_TOKEN",
  "<你的 Access Token>",
  "User"
)

3. Install the Plugin

Install the plugin into the DSH Web profile:

dsh plugin --profile web add github:hamliy-feng/duhai-vision
dsh web

Users on Linux and macOS should use the corresponding commands:

python3 -m pip install "paddleocr>=3.4,<4"
export PADDLEOCR_ACCESS_TOKEN="<你的 Access Token>"
dsh plugin --profile web add github:hamliy-feng/duhai-vision
dsh web

After installation, DSH obtains the Duhai Vision · deepseek-v4-flash model routing and can paste or upload images directly in the chat box.

Features

  • Structured output: Returns structured JSON, text content, layout, tables, formulas, etc.
  • Auditability: Provides model, page count, elapsed time, and token data, avoiding unknown values being recorded as 0.
  • Default preference: auto uses PaddleOCR-VL first for all supported vision tasks.
  • Fallback mechanism: Qwen is enabled only when explicitly selected or when PaddleOCR-VL fails.

Workflow

  1. Image processing: Chat-box images are first saved to a content-hash path under ~/.dsh/duhai-vision/attachments/.
  2. Structured extraction: PaddleOCR-VL extracts structured observations.
  3. Model inference: DeepSeek then performs inference and generates the answer.
  4. Text-only requests: Text-only requests are passed directly to DeepSeek.

Notes

  • Free quota: Each user currently has a free quota of 3,000 pages per model (subject to official rules).
  • Privacy and security: Remote services will receive images. Sensitive materials should be de-identified first, or avoid using the remote route.