Introduction¶
DeepSeek Harness (DSH) uses an “everything is a plugin” architecture. Plain-text models (for example, deepseek-v4-flash) cannot directly declare image input, which makes them unable to handle visual tasks directly. The @qizhen2021/dsh-plugin-vision plugin registers a see tool. It outputs a plain-text three-channel analysis for any image: offline OCR (with coordinates), ASCII layout maps, and semantic descriptions from a vision-language model (VLM). This enables any model (including plain-text models) to “see” images through the tool.
Core Features¶
The plugin provides three processing channels:
- OCR (offline): Uses the macOS Vision framework and supports automatic language detection. The output format is
x y w h|text, with coordinates as normalized percentages ×100, using the bottom-left as the Y-axis origin and sorted by line. - ASCII layout maps: Uses PIL to generate a grayscale image and converts it into an 88-column character art, with the gradient
' .:-=+*#%@'. Plain-text models can directly “see” the layout. - Semantic descriptions: Produces a structured report using the
mimo-v2.5model onopencode.ai(default). It supports injecting OCR results as ground truth to ensure verbatim string accuracy.
Installation and Enabling¶
The DSH GUI plugin manifest page is read-only and does not support local installation. It must be installed via Profile Bundles.
- Install the package:
cd ~/.dsh/profiles/web && pnpm add @qizhen2021/dsh-plugin-vision
- Modify
~/.dsh/profiles/web/package.jsonand append the following to thedsh.profile.bundlesarray:
{
"dsh": {
"profile": {
"bundles": [
"@deepseek-ai/dsh-base",
"@deepseek-ai/dsh-web-app",
"@qizhen2021/dsh-plugin-vision" // ← 追加这一行
]
}
}
}
- Restart
dsh webto activate the plugin.
Configuration¶
The plugin supports multiple configuration options, which can override the defaults in ~/.dsh/profiles/web/cordis.patch.yml.
| Key | Default | Description |
|---|---|---|
defaultModel |
mimo-v2.5 |
Default vision model, priced the same as flash |
gatewayBaseUrl |
https://opencode.ai/zen/go/v1 |
OpenAI-compatible gateway address |
credentialPath |
~/.dsh/.credentials.yaml |
Credential document path |
credentialKeys |
["OPENCODE_GO_API_KEY", "OPENCODE_API_KEY"] |
Candidate key names |
maxImageSide |
1600 |
Maximum image side for VLM preprocessing |
asciiWidth |
88 |
Number of columns in ASCII layout maps |
vlmMaxTokens |
1200 |
Maximum output tokens for the VLM |
ocrTimeoutMs |
90000 |
OCR timeout, including compilation |
vlmTimeoutMs |
120000 |
Total timeout for gateway calls |
Credentials use a two-layer reading strategy: they are preferentially read from the ctx.credentials service, and if that service is not mounted, the plugin falls back to reading the YAML file directly from credentialPath. Credentials are not in code, not in logs, and not in the schema.
Tool Usage¶
Call the see tool with the following parameters:
- Parameters:
file_path(required),ocr?/ascii?/vlm?(all default to true), andmodel?(defaults to the configureddefaultModel). - Return value: A structured report in JSON format, containing the results from the three channels. A failure in any channel does not cause the whole operation to fail; it is only marked as
ok: falsewith an error message. - Error handling:
- File-related errors return
FsError. - Vision-processing errors return
VisionError.
- File-related errors return
Notes and Limitations¶
- Platform limitation: OCR is supported only on macOS (Vision framework), while ASCII and VLM are cross-platform.
- Performance overhead: The first Swift call includes a compilation step, taking about 1–2 seconds; the timeout is set to 90s.
- No screenshot mode: The current version does not support screenshots.
- Credential security: Credential handling is entirely in-process and is not exposed in code or logs.
- Image processing: Large images must be resized via
prep.pyso that the longest side is ≤1600, otherwise gateway or memory issues may occur.