Text-based agents usually cannot directly understand visual information. SideSight provides a CLI-first vision sidecar for DeepSeek Harness, allowing agents to analyze screenshots, charts, UI diffs, and videos through OpenAI-compatible multimodal models. It supports filesystem skills and MCP servers (optional), and provides native offline OCR capability for macOS.
Core Features¶
- CLI-first vision sidecar: Uses the command-line interface as the interaction method to bridge vision capabilities to text agents.
- OpenAI-compatible multimodal models: Supports configuring multimodal models compatible with the OpenAI API format.
- Filesystem skill support: Can be integrated directly into project directories as a filesystem skill.
- MCP server support (optional): Provides optional MCP server integration.
- Local OCR support (macOS): Performs offline OCR using the native Vision framework on macOS.
Installation and Enablement¶
Install as a DSH plugin¶
In a DSH session, have the agent run the installation, or run the command directly:
npx @deepseek-ai/dsh plugin --profile web add github:ZhuXinAI/sidesight
After changing the configuration, restart the profile:
npx @deepseek-ai/dsh --profile web
Find the SideSight entry in the Web settings interface, and fill in Base URL, Vision model, and API key.
Install as a filesystem skill¶
DeepSeek Harness automatically discovers skills from .dsh/skills/ or $DSH_HOME/skills/.
- Create the skill directory at the project root:
mkdir -p .dsh/skills
- Clone the repository:
git clone https://github.com/ZhuXinAI/sidesight .dsh/skills/sidesight
Configuration and Usage¶
Model selection¶
After configuration is complete, before sending or pasting an image, you must select the DeepSeek (SideSight vision) provider. Ordinary text models (such as DeepSeek-V4-Flash or deepseek-official) do not support image input and will directly reject attachments.
API Key Security¶
The API key field is write-only. Configured keys are only displayed in the settings interface, are not echoed in chat records or ordinary tool results, and are not interpolated into command text.
Config file settings¶
You can also set the plugin fields directly in the cordis.patch.yml config file:
- update:
id: sidesight
config:
baseUrl: https://provider.example/v1
apiKey: your-key
model: your-vision-model
Typical Usage¶
Guide the agent to perform vision tasks through shell commands or instructions:
- Analyze screenshot causes:
Look at /path/to/screenshot.png and tell me why the button is disabled.
- Read error text:
Read the exact error text from /path/to/error.png.
- Compare UI changes:
Compare before.png and after.png and list only visible UI differences.
- Use local OCR:
Use native offline OCR on /path/to/receipt.png.
Notes¶
- MCP note: Neither the DSH plugin path nor the filesystem skill path starts an MCP server.
- Sensitive data: The API key is stored in the config file. Treat it as sensitive data.
- Permission scope: The plugin runs with the current DSH process permissions. Check the source code and license before installing.
- Local OCR limitations: The local OCR feature currently supports only macOS.