In the development practice of DeepSeek Harness (DSH), enabling models to understand images often stops at descriptive text. General-purpose vision-language models (VLMs) provide visual impressions rather than geometric facts; when directly asked about distances or positions, models are prone to giving incorrect answers or fabricating coordinates. dsh-tool-accurate-vision addresses this pain point by turning visual reasoning into computable structured data.

Feature Positioning

This is a model-facing accurate_vision tool. It processes image files through an OpenAI-compatible vision model and returns structured notes as well as bounding box primitives normalized to 0-1000. The tool formats these primitives into <vision-context> XML blocks for a pure text model to read in the next conversation turn, thereby obtaining accurate object location, layout, and OCR information without losing spatial fidelity.

The tool provides the following capabilities:
* Structured output: Returns structured notes and normalized bounding boxes (0-1000).
* XML context encapsulation: Packages primitives into <vision-context> blocks for direct parsing by text models.
* Visual feedback: Generates self-contained annotated SVG files for manual verification.
* Provider-agnostic: Supports any OpenAI-compatible multimodal endpoint.
* Auxiliary computation: Includes the built-in bboxEdgeDistance function for calculating edge distances.

Installation and Configuration

Installing the plugin requires DSH’s plugin manager.

dsh plugin --profile web add dsh-tool-accurate-vision

Because the vision API key is separated from DeepSeek’s main key, set the environment variable separately after installation:

export VISION_API_KEY=sk-...

After installation, override the configuration in cordis.patch.yml for the corresponding profile to adapt to your environment. Configuration items include the model name, endpoint URL, API key reference, whether to request primitives, whether to generate SVG, maximum tokens, timeout, temperature, and whether to disable thinking.

- id: tool-accurate-vision
  config:
    model: gpt-4o
    baseURL: https://api.openai.com/v1
    apiKeyEnv: VISION_API_KEY
    primitives: true
    annotate: true
    maxTokens: 8192
    timeoutSecs: 120
    temperature: 0
    disableThinking: true

Note: disableThinking: true skips the reasoning stage and is suitable for scenarios with high requirements for speed and stability.

How It Works

The invocation flow is as follows:

  1. Receives an image file and converts it into a Base64 Data URL.
  2. Sends it to the configured OpenAI-compatible vision model.
  3. The model returns structured notes and bounding-box primitives in JSON format.
  4. The tool encapsulates the primitives into a <vision-context> XML block for subsequent text models to read.
  5. The tool also generates an annotated SVG file containing all bounding boxes and labels.

Typical Application Scenarios

The core value of this tool lies in transforming visual impressions into geometric facts. In scenarios requiring precise distance calculations, it significantly improves the verifiability of results.

For example, given a hand-drawn physicist network diagram, if you ask “which node is physically closest to Marie Curie?”, a standard vision model can only provide an intuitive judgment (often incorrect). After using this tool, each node carries a verifiable 0-1000 bounding box. The agent can obtain exact data by calculating edge distances using the tool’s built-in bboxEdgeDistance function.

For instance, the calculation results show that Picard’s edge distance is 25.96, while Langevin’s is 58.00, leading to the correct conclusion. This structured data enables text models to perform rigorous geometric calculations rather than relying on unreliable intuition.

The plugin is released under the MIT License.