Foreword¶
In the DeepSeek Harness (DSH) workflow, an Agent’s claim of completion (“Done”) is often only a subjective judgment, potentially hiding failed tests, build errors, or missed requirements. To address this issue, an independent acceptance layer is needed to verify the Agent’s output. This article introduces the dsh-autopilot plugin.
What It Is¶
dsh-autopilot is an acceptance-driven auto-completion engine for DeepSeek Harness. The project is maintained by 245678000000 and is available under the MIT open-source license. It addresses the problem that an Agent’s self-reported completion does not mean the task has truly been completed. By performing independent acceptance checks with deterministic evaluators, it ensures that a completion certificate is issued only when all required checks pass.
Core Features¶
The plugin mainly provides the following capabilities:
- Independent acceptance and auto-completion: It hooks into the official
agent/turn-stoppingwithout replacingctx.goals, and independently evaluates the Agent’s output. - Repair loops and retry mechanisms: For fixable failures, it generates structured feedback and continues the current turn via
agent.steer, or hands off to the official driver for the next round. - Certificate generation: After passing, it writes out a
CompletionCertificateand explicitly marks the verification status. - Multiple evaluators: It supports various evaluation methods, including command (command execution), file (file checks), regex (regular expression matching), git (workspace state), http (HTTP requests), manual (human confirmation), and agent (independent judging).
- Regression / no-progress / oscillation detection: It compares results with the previous round to prevent the model from spinning its wheels in the wrong direction or oscillating between A and B.
- Blocking mechanisms: It blocks situations such as missing keys, denied approvals, or required criteria that cannot be verified, avoiding ineffective retries.
Installation and Activation¶
The system requires Node.js 20.19+ and an installed DeepSeek Harness.
Install it as a Harness plugin:
dsh plugin add https://github.com/245678000000/dsh-autopilot
Or use a local checkout as an overlay:
pnpm dsh web --patch /absolute/path/to/dsh-autopilot/cordis.patch.yml
Typical Usage¶
The plugin provides a standalone CLI tool that can evaluate a project without running the full Harness.
Standalone evaluation (without executing shell commands):
dsh-autopilot evaluate --cwd ./fixtures/file-only
Standalone evaluation (with authorization to execute shell commands):
dsh-autopilot evaluate --cwd ./examples/login --allow-commands
Common commands also include viewing status, criteria configuration, and certificates:
dsh-autopilot status
dsh-autopilot criteria
dsh-autopilot certificate
Run demo scenarios:
npm run demo
Applicable Scenarios and Caveats¶
dsh-autopilot is suitable for workflows that require strict acceptance criteria and need to prevent Agent self-deception.
Caveats include:
- Untrusted configuration:
.autopilot.ymlis treated as untrusted project configuration; broken acceptance criteria can result in an incorrect certificate being issued. - Probabilistic evaluation: LLM evaluators are probabilistic; invalid output is recorded as
error, notpass. - System availability: External systems may be unavailable.
- Mathematical proof limitations: Passing tests mathematically cannot prove that the software is correct.
- Permission risks: The plugin runs with the current DSH process permissions; review the source code and license before installation.
- In-process call limitations: If a trusted in-process plugin directly calls
ctx.goals.complete(), the official Goal service still acknowledges it. Autopilot intercepts the model tool calls, not all in-process calls.
Summary¶
By introducing an independent acceptance layer, dsh-autopilot transforms the definition of “Done” from a subjective declaration to objective verification, improving the reliability of Harness workflows. For more details, refer to the GitHub repository.