Introduction

In DeepSeek Harness (DSH) workflows, verifying plugin changes or preset differences is a common need. dsh-plugin-compare provides this capability. It turns execution evidence into a comparison timeline and reports.

Positioning

  • Name: dsh-plugin-compare
  • Author: yminghua
  • Purpose: Controlled A/B comparison and evidence-based reporting for DSH plugins and presets.

Core Capabilities

  1. Compare plugins and presets: Supports comparison based on existing sessions or running controlled A/B experiment groups.
  2. Evidence reports: Generates side-by-side timelines, measurable deltas, and explicit success-check results.
  3. Controlled runs: Automatically duplicates the workspace, submits the same prompts, and captures Git evidence.
  4. Export formats: Supports multiple formats, including HTML, JSON, SVG, PNG, and more.
  5. Statistical metrics: Displays mean and median deltas for time and tokens, along with Student-t 95% intervals.

Installation and Enablement

This plugin requires Node.js 22.19+ and a DSH Web profile.

Install the stable version:

dsh plugin --profile web add dsh-plugin-compare

Install the prerelease version (requires explicit declaration):

dsh plugin --profile web add dsh-plugin-compare@next

After installation or update, stop the currently running Web Host and restart it:

# 停止运行
Ctrl+C
# 启动
dsh web

After restarting, a Compare panel appears in the DSH Web interface.

Typical Usage

Prepare the test environment:

node example/prepare.mjs

Then you can follow the Chinese step-by-step guide in the repository to compare presets such as Standard and dsh-expert-mode, check the results, and export reports.

Notes

  1. Data privacy: Exported HTML, JSON, SVG, PNG, and screenshots may contain prompts, source paths, code diffs, command output, provider/model names, and session IDs. Automatic redaction is best-effort and is not privacy approval. You must manually review and remove sensitive information before sharing.
  2. Statistical meaning: The intervals describe observed variation with small samples and do not automatically establish a winner or causal conclusion.
  3. File handling: Dependency directories are excluded from copies, external symbolic links are rejected, and Git worktree pointer files are not copied back into the experimental environment.
  4. Failure handling: If the Agent fails before completing the task, structured failure information is displayed and exported, the success check command is marked as not run, and the comparison is marked invalid.

References

  • GitHub repository: https://github.com/yminghua/dsh-plugin-compare
  • Skill Center directory: https://www.skillhub.cn/plugins/yminghua/dsh-plugin-compare