Introduction

DeepSeek Harness (DSH) follows the “everything is a plugin” philosophy. When building agents based on text-only models, visual interaction capabilities, such as reading images and generating images, are often a weak spot. The dsh-vision plugin injects image recognition and generation capabilities into the text model by taking over the input event stream and extending the toolchain, enabling it to describe local images or generate visual content.

Plugin Positioning

This is a DSH plugin maintained by HarryLi-7. It mainly solves the problem that text models cannot directly process image inputs and outputs, supports a multi-engine failover chain, and provides complete UI integration.

Installation and Activation

Add the plugin via the command line:

dsh plugin --profile web add /path/to/dsh-vision

After installation is complete, restart the dsh web service.

Core Features

Image Recognition

Use the describe_image tool to read local images.
* Output modes: By default, it returns structured JSON, including description, OCR-recognized lines, layout regions, and entity information. It can also be configured for plain-text mode.
* Engine chain: Codex → Gemini → Zhipu GLM.

Image Generation

Use the generate_image tool to generate images from prompts.
* Quality options:
* auto: Nano Banana 2 Lite (economy mode).
* hd: Nano Banana Pro 1K/2K.
* 4k: Nano Banana Pro 4K.
* Display: Generated results are displayed as cards in the conversation, supporting preview, download, and navigation to the file system.

Input Methods and File Processing

The plugin intercepts paste and drag-and-drop operations at the window level.
* Upload path: Images are uploaded to ~/.dsh/generated-images/uploads/, using content-addressed storage, so identical images are retained only once.
* Model interaction: The actual content sent to the model is the file path as text, and the model processes text only.
* UI feedback: A thumbnail preview bar is provided above the input box, supporting horizontal scrolling and single-image deletion, with the corresponding path removed synchronously.

Storage Strategy

The plugin follows a strict temporary-storage principle.
* Codex sessions: It uses the --ephemeral and --sandbox workspace-write parameters and does not generate persistent session files.
* Uploaded files: Files are retained for 7 days by default and can be configured through uploadRetentionDays in Settings. Setting it to 0 retains them permanently. Orphaned .meta.json sidecar files are cleaned up automatically.

Configuration and Credentials

Credential Configuration

Configure the key for each engine in ~/.dsh/.credentials.yaml:
* DEEPSEEK_API_KEY: DeepSeek primary engine.
* GEMINI_API_KEY: Gemini recognition (free tier; Pro will be downgraded to Flash).
* GEMINI_IMAGE_API_KEY: Gemini generation (requires payment and belongs to a separate project).
* ZHIPU_API_KEY: Zhipu AI fallback engine.

Engine Parameter Settings

Find the dsh-vision configuration entry on the DSH Settings page:
* codexPath: Codex CLI path (leave empty for automatic detection).
* uploadRetentionDays: Number of days to retain uploaded files.
* describeEngines / generateEngines: Control the engine invocation order for recognition and generation.
* Engine model parameters: For example, Codex can be configured with parameters such as gpt-5.6-luna, max, and priority.

Environment Requirements

  • A Node.js environment with the DSH harness (dsh web) installed.
  • Codex CLI, installed via npm install -g @openai/codex, and a logged-in ChatGPT account.
  • Optional: A Zhipu or Gemini key for fallback engines.

Summary

dsh-vision provides DSH with reliable visual processing capabilities through multi-engine failover and an ephemeral storage policy. Developers can adjust the engine chain and parameters in Settings as needed to meet different recognition and generation requirements.

Plugin Directory
GitHub Repository