Introduction

DeepSeek Harness (DSH) is a plugin-driven agent runtime environment. When building agents that can handle local files, it is common to encounter scenarios where text needs to be extracted from images. If an external OCR API is used, it not only requires network requests but also involves key management. The dsh-macos-vision-ocr plugin leverages Apple’s Vision framework to provide offline OCR capability locally, allowing image content to be read without network requests.

Plugin Introduction

Name: dsh-macos-vision-ocr
Core Value: An offline text recognition tool based on the macOS Vision framework, providing DSH with the ability to read images.
Maintainer: leozou320-ai
License: MIT

Core Features

This plugin provides the ocr_image tool, with the following characteristics:

  1. Local Recognition: Uses VNRecognizeTextRequest for text recognition, with high default accuracy.
  2. Format Support: Supports common image formats, including PNG, JPEG, WebP, GIF, TIFF, BMP, HEIC, and HEIF.
  3. Multilingual: Supports specifying BCP-47 recognition language codes for each call.
  4. Secure Execution: Uses a fixed subprocess argument vector; image paths are not directly concatenated into shell commands, preventing injection risks.
  5. Caching Mechanism: Compiles the embedded Swift helper on first use, then reuses it through content-addressed caching to avoid repeated compilation.
  6. Output Control: Returns bounded output results with a truncated flag indicating whether the recognized text has been truncated.

Installation and Activation

Installing a plugin in DSH requires specifying the target profile:

dsh plugin --profile web add github:leozou320-ai/dsh-macos-vision-ocr

After installation, the profile must be restarted for the changes to take effect.

Typical Usage

There are two ways to use this tool:

  1. Explicit Invocation: Pass parameters in a configured JSON tool call.
  2. Agent Interaction: Directly instruct the Agent to “read this image.”

Example of explicit invocation:

{
  "file_path": "./scan.png",
  "languages": ["zh-Hans", "en-US"]
}

Notes and Limitations

  1. Environment Requirements: Must run on macOS 13 or later, with Xcode Command Line Tools installed (the swiftc command must be in PATH), and DeepSeek Harness version 0.1.0-rc.5 or later.
  2. System Compatibility: Because it depends on Apple Vision, this plugin is only available on macOS. It will fail on Linux and Windows due to missing underlying frameworks.
  3. First-Call Latency: The first tool call is slow because the Swift helper must be compiled; subsequent calls read from the cache and are faster.
  4. Privacy and Security:
    • OCR runs entirely locally and makes no network requests.
    • Image paths are checked by the Harness file system service before the native helper is executed.
    • The recognized text becomes tool output, entering the current session log and model context.
  5. Functional Boundaries: This plugin only extracts text and does not have general visual understanding capabilities (such as recognizing objects, faces, or scenes).
  6. Deployment Recommendation: Before deployment in sensitive environments, review the source code of third-party plugins and pin versions.