Introduction¶
When developing agents that need to operate a desktop, enabling the model to directly understand screen content often relies on multimodal models. This article introduces the OmniVision plugin, which uses OmniParser to convert the screen into structured data, allowing the model to “see” and interact with the interface without native multimodal vision capabilities.
Overview¶
OmniVision is a GUI agent plugin based on OmniParser and maintained by xiaozhengdeng. It converts a desktop or arbitrary images into structured elements (text + icons + pixel coordinates), addressing the difficulty of GUI interaction when the model lacks multimodal capabilities.
Core Features¶
The plugin provides the following core capabilities:
* Screen recognition: Captures the desktop or parses images and extracts interactive elements with pixel coordinates using OmniParser.
* Desktop automation: Supports click, double-click, right-click, drag, input, key presses, hotkeys, and mouse wheel operations, executable by element ID or raw coordinates.
* Real-time recognition Dock: Provides a real-time recognition view with SOM annotations, supporting hover highlighting and click-to-zoom.
* Recognition history: Records previous recognition results and shows additions, deletions, and changes compared with the latest result.
* One-click summary: Sends the currently recognized structured information to the session and asks the model to generate a summary.
* Image parsing: Selects local image files through the Dock for parsing, bypassing the model’s multimodal limitations.
* Call history: Tracks each gui_* tool invocation and the corresponding action execution status.
Installation and Activation¶
The installation command is as follows:
dsh plugin --profile web add dsh_omnivision
After installation, you must restart the web process to load the plugin. The plugin is loaded as a profile bundle layer and is divided into a Host half (tool registration) and a Client half (browser interface).
Prerequisites¶
The following conditions must be met before use:
* Platform: Windows only.
* OmniParser service: Ensure the OmniParser FastAPI service is running on 127.0.0.1:8000.
* Python dependency: The Python venv in the runtime environment must have pyautogui installed for screenshot capture and input automation.
Tool Reference¶
The plugin registers the following gui_* tools in the shared tool registry:
* gui_capture: Captures the screen and runs OmniParser to extract interactive elements, refreshing the visual state.
* gui_act: Performs actual mouse/keyboard operations on the desktop.
* gui_find: Searches elements from the last recognition based on text or type.
* gui_state: Views the current visual state without re-parsing.
* gui_verify: Recaptures and re-parses the screen to repeatedly confirm the presence or absence of specific text.
* gui_task: Executes a multi-step UI plan, supporting re-parsing and assertions between steps.
* gui_open_app: Starts an installed desktop application via its Windows AUMID.
* gui_parse_image: Parses third-party images (attachments) in the session into the shared visual state.
Usage¶
The model side directly invokes the gui_* tools above to interact with the plugin. The OmniVision Dock on the browser side provides a visual operations panel, including a smart recognition view, image parsing, one-click summary, recognition history, and call history.
Notes¶
- The plugin runs with the permissions of the current dsh process; ensure you have permission to operate the target application.
- Be sure to review the source code and license before use to ensure compliance with usage norms.
Conclusion¶
OmniVision provides a standardized GUI operation interface for DeepSeek Harness, making it suitable for agent development scenarios that require handling desktop tasks. For more details, see the official directory or the GitHub repository.