Introduction

DeepSeek Harness (DSH) uses an “everything is a plugin” architecture. Plain-text models (for example, deepseek-v4-flash) cannot directly declare image input, which makes them unable to handle visual tasks directly. The @qizhen2021/dsh-plugin-vision plugin registers a see tool. It outputs a plain-text three-channel analysis for any image: offline OCR (with coordinates), ASCII layout maps, and semantic descriptions from a vision-language model (VLM). This enables any model (including plain-text models) to “see” images through the tool.

Core Features

The plugin provides three processing channels:

  1. OCR (offline): Uses the macOS Vision framework and supports automatic language detection. The output format is x y w h|text, with coordinates as normalized percentages ×100, using the bottom-left as the Y-axis origin and sorted by line.
  2. ASCII layout maps: Uses PIL to generate a grayscale image and converts it into an 88-column character art, with the gradient ' .:-=+*#%@'. Plain-text models can directly “see” the layout.
  3. Semantic descriptions: Produces a structured report using the mimo-v2.5 model on opencode.ai (default). It supports injecting OCR results as ground truth to ensure verbatim string accuracy.

Installation and Enabling

The DSH GUI plugin manifest page is read-only and does not support local installation. It must be installed via Profile Bundles.

  1. Install the package:
    cd ~/.dsh/profiles/web && pnpm add @qizhen2021/dsh-plugin-vision
  1. Modify ~/.dsh/profiles/web/package.json and append the following to the dsh.profile.bundles array:
    {
      "dsh": {
        "profile": {
          "bundles": [
            "@deepseek-ai/dsh-base",
            "@deepseek-ai/dsh-web-app",
            "@qizhen2021/dsh-plugin-vision"   // ← 追加这一行
          ]
        }
      }
    }
  1. Restart dsh web to activate the plugin.

Configuration

The plugin supports multiple configuration options, which can override the defaults in ~/.dsh/profiles/web/cordis.patch.yml.

Key Default Description
defaultModel mimo-v2.5 Default vision model, priced the same as flash
gatewayBaseUrl https://opencode.ai/zen/go/v1 OpenAI-compatible gateway address
credentialPath ~/.dsh/.credentials.yaml Credential document path
credentialKeys ["OPENCODE_GO_API_KEY", "OPENCODE_API_KEY"] Candidate key names
maxImageSide 1600 Maximum image side for VLM preprocessing
asciiWidth 88 Number of columns in ASCII layout maps
vlmMaxTokens 1200 Maximum output tokens for the VLM
ocrTimeoutMs 90000 OCR timeout, including compilation
vlmTimeoutMs 120000 Total timeout for gateway calls

Credentials use a two-layer reading strategy: they are preferentially read from the ctx.credentials service, and if that service is not mounted, the plugin falls back to reading the YAML file directly from credentialPath. Credentials are not in code, not in logs, and not in the schema.

Tool Usage

Call the see tool with the following parameters:

  • Parameters: file_path (required), ocr?/ascii?/vlm? (all default to true), and model? (defaults to the configured defaultModel).
  • Return value: A structured report in JSON format, containing the results from the three channels. A failure in any channel does not cause the whole operation to fail; it is only marked as ok: false with an error message.
  • Error handling:
    • File-related errors return FsError.
    • Vision-processing errors return VisionError.

Notes and Limitations

  • Platform limitation: OCR is supported only on macOS (Vision framework), while ASCII and VLM are cross-platform.
  • Performance overhead: The first Swift call includes a compilation step, taking about 1–2 seconds; the timeout is set to 90s.
  • No screenshot mode: The current version does not support screenshots.
  • Credential security: Credential handling is entirely in-process and is not exposed in code or logs.
  • Image processing: Large images must be resized via prep.py so that the longest side is ≤1600, otherwise gateway or memory issues may occur.