DeepSeek Harness (DSH) adopts a plugin-based architecture, where the capabilities of the Agent depend on the installed plugins. When developing tasks involving desktop operations, the Agent can often only “see” text and cannot directly interact with the Windows interface. The Vision Use plugin solves this pain point by connecting to the vision channel and system-level operation interfaces.
What It Is¶
Vision Use is a DeepSeek Harness plugin maintained by user zzy6-a and released under the MIT license. It primarily addresses the problem that the Agent cannot directly perceive the desktop environment or perform mouse and keyboard operations.
Core Features¶
The plugin provides the following tool interfaces:
view_screen: Captures the full Windows screen and sends it directly to the model’s vision channel (not OCR—it really views the image).view_image: Sends any image file to the vision channel (supports Linux / Windows paths).computer_move: Smoothly moves the cursor to the specified coordinates (x, y).computer_click: Performs left, right, or double-click operations.computer_type: Types text. By default, it uses real keystroke injection key by key (zero clipboard); Chinese input goes through the input method flow.computer_key: Sends key combinations (such asctrl+t,enter,alt+f4).computer_overlay: Starts or stops the operation overlay, with optional automatic shutdown after idle time.computer_status: Queries overlay status, ESC flag, current mode, and cursor position.
Installation and Enabling¶
Use the official installation command to add the plugin:
dsh plugin --profile web add https://github.com/zzy6-a/vision-use/releases/download/v0.2.0/dsh-vision-0.2.0.tgz
After installation, restart the DSH client, refresh the browser, and confirm that the status capsule appears below the input box.
Typical Usage¶
After enabling the plugin, the Agent can perform the following actions through natural-language commands:
- Look at my screen
- Search for the DeepSeek official website on Bing and open it
- Open Notepad and type some text
The Agent will automatically start the overlay, perform a view_screen visual inspection, invoke tools from the computer_* series to carry out operations, and take another screenshot to verify when necessary.
Use Cases and Considerations¶
System Environment¶
- Supported environments: Native Windows and WSL2 + Windows. The plugin automatically detects the host environment and selects the corresponding path/subprocess approach.
- Unsupported environments: Linux / macOS hosts are not currently supported; only native Windows and WSL + Windows are supported. On other systems, desktop-control tools will report explicit errors, while
view_imageremains available.
Compatibility Limitations¶
- Input method: Chinese / non-ASCII characters go through the input method flow, with Pinyin letters and spaces sent to commit the text.
- Chromium: Chromium browsers ignore
KEYEVENTF_UNICODEinjection, so real keystroke injection is required. - WinUI: WinUI applications such as Notepad do not recognize injected Ctrl key combinations.
- Custom-drawn UIs: In applications with custom-drawn UIs (such as the Edge tab bar), coordinates cannot be trusted; visual localization is recommended first.
- Cursor overriding: Software that renders custom cursors (such as some games or Electron apps) cannot have its cursor replaced by the system-level cursor.
Privacy¶
Screenshots are added as image attachments to the current session context and are ultimately sent with the session request to the selected model. The plugin itself does not make network requests or call any APIs.
Conclusion¶
Vision Use combines visual perception with system-level operations, filling a gap in desktop interaction for DSH Agents. After installation and enabling, the Agent can understand screen content through the vision channel and use mouse and keyboard tools to perform tasks. For project details and source code, see the GitHub repository.