Introduction

In model inference scenarios, enabling an Agent to see the screen usually requires manual intervention—manually taking a screenshot, pasting it, and then asking the model to analyze it. This breaks the continuity of the automated process.

The dsh-screen-reader plugin solves this problem. It allows an Agent to directly capture screen images during execution and pass the image itself as a content block to the current model, rather than having it relayed by a second model. In addition, it provides local exact pixel diff capabilities: local computation determines whether anything changed and where it changed, while the model only answers what it changed into.

Core Features

The plugin provides two main tools: see_screen and see_diff.

see_screen: Capture the Screen and Pass It Directly to the Model

This tool captures the screen, a window, or a specific region, and provides the image directly to the current model for viewing, without going through any intermediate layer.

  • Parameters:
    • window: Capture a specific window.
    • region: Capture a specific region.
  • Image processing: Images are forcibly normalized on the provider side, with an upper limit of 384 visual tokens.

see_diff: Local Pixel Diff

This tool compares two images, uses local exact pixel diff to locate changed regions, and then hands the changed regions to the model for explanation.

  • Workflow:
    1. Compute the differences locally.
    2. Locate the changed regions.
    3. The model only answers what it changed into, and is not responsible for determining whether it changed.

Installation and Enabling

The plugin only supports Windows.

  1. Use the following command to install the plugin (via the profile bundle method):
    dsh plugin --profile <profile-name> add <package-name-or-local-path>
For example, install from a local directory:
    dsh plugin --profile <profile-name> add C:\path\to\dsh-screen-reader
  1. After installation, restart DSH to activate the new plugin.

Use Cases and Considerations

Use Cases

  • The Agent needs to self-verify the results of operations during execution (for example: checking that a setting change was applied correctly).
  • The changed regions between two images need to be identified.

Notes and Limitations

  • Platform limitation: Only supports Windows, and depends on PowerShell 5.1’s System.Drawing and PrintWindow.
  • Visual capability: The plugin does not enhance the model’s visual capability; it only passes images to the model.
  • Measurement precision: The pixel error for precise measurements can reach ±40%, so it is primarily for description rather than measurement.
  • Recognition challenges: Low-contrast differences are a danger zone, and shape semantics are a weakness.
  • Resolution and cropping: Resolution does not affect accuracy; cropping is a way to increase detail density.
  • Time and causality: Only one frame is available, without temporal or causal context (the screen scrolling memory feature has been removed).
  • Storage management: The attachment library will grow, and there is no cleanup mechanism.

Appendix

  • Directory page: https://www.skillhub.cn/plugins/cbg33695/dsh-screen-reader
  • GitHub: https://github.com/cbg33695/dsh-screen-reader
  • License: MIT