DeepSeek Harness (DSH) adopts an “everything is a plugin” architecture, which makes extending its functionality very straightforward. In agent development, a common requirement is to compare different models on the same code task, including tool-call trajectories, code change volume, and final response quality.

dsh-dual-model-eval plugin was designed for exactly this. It allows you to enable “Comparison test” mode in DSH’s model selector and run multiple configured model routes concurrently. Each model works independently in an isolated Git Worktree, with expandable real-time tool-call trajectories, line change statistics, and the ability to adopt the selected result as the baseline.

Core Features

This plugin primarily addresses concurrent multi-model evaluation and adoption:

  • Comparison test mode: Adds a switch to the existing model selector; after enabling, it supports concurrent comparison.
  • Multi-selection concurrency: Supports testing 2 to 4 configured model routes simultaneously.
  • Isolated execution: From the same shared Git commit point, it creates an independent Worktree for each model, ensuring they do not interfere with one another.
  • Real-time trajectories: Each model’s tool-call trajectory can be expanded independently and supports streaming viewing.
  • Statistical summary: Provides compact duration and tool-call count statistics; when expanded, it displays detailed metrics such as tokens, cache, TTFT, and more.
  • Code comparison: While displaying the final response, it also shows line change statistics (+added, -deleted, percentage) and the number of changed files.
  • File preview and download: Supports previewing a single modified file or downloading the complete candidate Worktree.
  • Result adoption: Provides an “Adopt this result” action, commits the selected patch locally, and fast-forwards the current workspace as the baseline for the next round.

Installation and Enablement

Before installing, please ensure your DeepSeek Harness version is 0.1.0-rc.7. This plugin needs to be installed into the built-in web profile:

npx @deepseek-ai/dsh@0.1.0-rc.7 plugin --profile web add github:huangdaxianer/dsh-dual-model-eval#v0.1.1

After installation, you need to restart the Web process for the changes to take effect:

npx @deepseek-ai/dsh@0.1.0-rc.7 web

Typical Usage

  1. Open the standard model selection menu and enable the Comparison test switch.
  2. Select at least two model routes that are already configured in Harness.
  3. Submit a coding request. Result cards appear immediately and display streaming progress while the models are running.
  4. Inspect each response and its expandable tool-call trajectory.
  5. Select Adopt this result under your preferred candidate. The next round of comparison will start from that commit’s code baseline and inherit the accepted conversation context.

Workspace and Git Behavior

  • Permissions: The plugin runs with local user permissions, not system-level permissions.
  • Worktree management: Each candidate model creates a new detached Worktree from the same base commit point. After the run finishes, the temporary Worktree is removed by default, while private evidence data is retained under ~/.dsh/dual-model-eval/.
  • Commits and pushes: The plugin can create local workspaces and commits, but it does not push to remote repositories and does not modify the global Git identity.
  • Non-Git workspaces: If the workspace does not have an initialized Git repository, the plugin automatically initializes one and stages files according to .gitignore.
  • Adoption workflow: The adoption action reconstructs the selected complete patch, creates a local commit with a command-scoped identity, and fast-forwards the source workspace.

Security and Notes

Because the plugin runs with local user permissions, it has the ability to create Worktrees and local commits within the workspace. Before installing third-party plugins, be sure to inspect the source code and license.

This plugin is suitable for scenarios where you need to verify behavioral differences between different models in real code environments, or where you need to quickly integrate the best model results into your project. By using local commits instead of remote pushes, developers can carefully review generated patches and evidence before publishing.