Introduction

In model inference scenarios, enabling an agent to “see” the interface is a common requirement. The traditional approach is for developers to manually take and upload screenshots. The DeepSeek Harness plugin dsh-screen-eye automates this process: the model can capture the screen while invoking tools and receive image content in the same call.

Core Features

This plugin provides two directly callable tools for DSH:

  • screenshot: Captures the screen and returns the image as a content block, which the model can view directly.
  • screen_permission: Reports permission status on macOS; when using action: "guide", it opens System Settings and navigates to the specific permission entry to guide the user in enabling it.

Installation

Installing this plugin requires no local build steps and has no additional dependencies.

dsh plugin --profile web add github:davidekingsss/dsh-screen-eye

After installation, restart DeepSeek Harness.

Permission Configuration (macOS)

macOS has strict permission controls for screen recording. This is the main obstacle to using this plugin.

Permission Status Detection

The plugin detects permission status by actually attempting to capture, rather than guessing. If permission has not been granted, screenshot fails. At this point, call the screen_permission tool; if action is "guide", it opens System Settings and guides the user.

Launcher Name Differences

The name of the macOS permission entry depends on the application that launched Harness (the “responsible process”), not Harness itself. For example:

  • Launched via a browser: the entry name is Google Chrome.
  • Launched via a terminal: the entry name is Terminal or iTerm.
  • Launched via a shortcut: the entry name is Run DeepSeek Harness.

The tool lists common cases, but users must manually match this entry in Settings.

Operation Steps

On first use on macOS, follow these steps:

  1. Call screenshot.
  2. If it fails, call screen_permission and set action: "guide".
  3. The system opens the Privacy and Security settings page.
  4. Find the corresponding entry in the Screen Recording list and toggle it on.
  5. Call screenshot again.

Permissions cannot be enabled automatically via programming; this is a limitation of the system security policy (SIP protects the TCC database).

Capture Modes

The screenshot tool supports multiple mode parameters, used to define the capture scope:

  • screen (default): Captures the main display.
  • display: Captures a single display by index.
  • region: Captures a rectangular area. The coordinate system origin is the top-left corner of the main display, so displays to the left of or above the main display may have negative coordinates.
  • displays: Captures no image; it only lists connected displays and their indices.
  • window: Captures the currently topmost window.
  • select: Interactive selection; waits for the user to click a window or drag a rectangle. This mode is not supported on Windows.

Multiple Frames and Motion

Set the frames parameter to greater than 1 to capture a continuous sequence of images. The interval_ms parameter determines the time interval between frames.

  • Sampling speed depends on the captured area size: a 4K full screen takes about 155 ms, and a 1200x800 area takes about 56 ms.
  • The returned images are stored as single frames and are not automatically composited into a GIF.
  • The total duration of the sequence is determined by timeout_ms (default 5 minutes), and the maximum frame count is limited by the provider (about 600 images).
  • The response reports the actual interval achieved.

System Differences (Windows)

The implementation on Windows differs significantly from macOS:

  • No permission reporting: Windows has no mechanism that allows the plugin to report screen recording permission status.
  • Image quality:
    • If the process is DPI-unaware, the captured image may be a low-resolution copy.
    • If the process is not attached to an interactive desktop (a non-interactive desktop), the capture may be black frames.

Limitations

  1. Single-call limit: Each screenshot call returns at most one image (even if the system has multiple displays).
  2. Dependency: The plugin itself has no native build steps, but on macOS it compiles a small helper from source on first capture.
  3. Manual intervention: macOS permissions must be enabled manually and cannot be automated.

Reference Links