Introduction

The core task of DeepSeek Harness (DSH) is to evaluate the combined effect of a model and its Harness configuration. In practical use, the same model (for example, V4 Pro) may exhibit completely different reasoning paths and execution trajectories on the same task because the Harness constructs different Request Surfaces. For example, it may sometimes appear as “Let me check…,” and sometimes as “We need to locate…”.

This difference is not simple model capability fluctuation, but the result of interaction between the model and the Harness. Existing evaluation methods often find it difficult to reproduce and compare a specific “Request Surface” while preserving context integrity. dsh-replay-lab aims to solve this problem by providing a controlled Request Surface replay and A/B experimentation workbench.

Plugin Positioning

dsh-replay-lab is a DeepSeek Harness plugin for replaying Request Surfaces and running A/B experiments on completed agent turns.

  • Core value: Freeze a turn, isolate candidate sessions, and compare Request Surfaces, trajectories, and costs.
  • Maintainer: tbxy09
  • License: MIT

It is mainly used to reproduce and debug long trajectories, repeated tool-call loops, no-progress turns, and regressions that depend on presets or plugins.

Core Features

This plugin focuses on the “Request Surface” rather than prompts alone, and provides the following capabilities:

  • Request Surface replay: Replay the full request context, not just the prompt. Supports preset or plugin candidates such as Standard, Minimal, and Anchored.
  • Turn freezing and isolation: Freeze a completed turn and isolate a candidate session for testing, while the source session and workspace remain unchanged.
  • Multi-dimensional comparison: Compare Request Surface differences, trajectory differences, and execution cost differences.
  • Sandbox evidence dashboard: Display evidence in an isolated sandbox environment. Supports custom prompts rather than only preset chips.
  • Request Surface Diff: Compare routing, stages, tool lists, and hash values.
  • Execution Delta: Show differences in token count, duration, steps, and tool invocations.
  • Evidence narrative generation: Automatically generate a fact-based comparative narrative.

Usage

Before using this plugin, ensure that you are running a DSH environment that meets the requirements.

Dependencies and Version

According to the package.json information, running this plugin requires Node.js version ^22.19.0 || >=24.

Code block: key information in package.json

{
  "name": "@webwalkerhq/dsh-replay-lab",
  "version": "0.1.6",
  "license": "MIT",
  "dependencies": {
    "@deepseek-ai/schemastery": "^3.18.1"
  },
  "peerDependencies": {
    "@deepseek-ai/cordis": "^4.0.1"
  }
}

Typical Usage

The typical workflow for using dsh-replay-lab is: freeze a turn, send a prompt, and obtain comparison results.

  1. Freeze the turn: Complete a turn in DSH and freeze it using the plugin.
  2. Isolate a candidate: Under the frozen state, select a candidate preset (such as Standard) or plugin for isolated execution.
  3. Open the sandbox dashboard: After isolated execution completes, open the sandbox evidence dashboard.
  4. Send a custom prompt:
    • In the dashboard, you can select a preset chip or directly enter any prompt.
    • Click Send.
    • The model will return an HTML-rendered chart or data.
  5. Request visualization:
    • Request Surface Diff: View similarities and differences in routing and tool lists.
    • Execution Delta: View changes in token count, duration, and tool invocations.
    • Evidence narrative: Get a fact-based comparative summary.

Notes and Use Cases

  • Data safety: The source session and workspace are never rewritten or rolled back during the entire process.
  • Temporary modifications: Changes made to candidate files during the run are restored to the replay checkpoint after completion and do not affect the source session.
  • Narrative generation mechanism: The evidence narrative is generated by an explicit, single direct model-run invocation, not by launching an Agent.
  • HTML compatibility: If the generated HTML is invalid, the dashboard falls back to the host chart.

This plugin is suitable for developers who need to deeply analyze the “model × Harness” combined effect in DSH, especially when precise reproduction of a specific Request Surface is required to diagnose regressions.

Conclusion

dsh-replay-lab provides a structured method to isolate the Request Surface from execution outcomes for independent observation. Through the sandbox dashboard and custom prompt capability, it makes long-trajectory debugging and cost analysis concrete and actionable.