Introduction

In DeepSeek Harness (DSH), when building agents, task outputs often depend on LLM sampling results, making “false completion” or “unfinished but self-reported completion” difficult to avoid. Existing DSH workflow plugins typically leave independent verification to the upper-layer architecture, leaving a runtime safety gap.

The deepseek-harness-reliability-governor plugin aims to fill this general runtime gap. It does not try to make the LLM deterministic; instead, through a “reliability contract,” during contract activation, it steers agent behavior until an observable check passes, bounded repair exhaustion occurs, or a truthful report is attested. It turns the model’s one-off assertions into an evidence-based deterministic decision process.

Core Capabilities

The plugin provides a set of commands for managing evidence gating and code verification workflows:

  1. reliability_assess: Before evaluating task output, preview declarative coverage scope, independent source counts, and fragile evidence warnings.
  2. reliability_draft: In optional auxiliary model mode, request a restricted text-only draft and record its source and receipt.
  3. reliability_begin: Start an explicit completion contract.
  4. reliability_begin_code: Start a code contract that automatically includes trusted verification profiles required for each deployment.
  5. reliability_verify: Immediately run deterministic checks.
  6. reliability_status: Read the persisted contract, attempt count, terminal state, and receipt.
  7. reliability_abstain: Stop without fabricating proof.
  8. reliability_code_profiles: List metadata for trusted profiles without exposing model-rewritable commands.
  9. reliability_code_verify: Execute immutable profiles through Harness-managed subprocess and sandbox services.
  10. agent/turn-stopping enforcement: Validate active contracts before allowing settlement, then guide bounded repair or truthful report attestation/exhaustion.

Review Mechanism and Intent Confirmation

By default, reliability_begin pauses twice before activation. The first pause uses the A2UI v0.9.1 base catalog interface to present the interpreted objective, constraints, assumptions, non-goals, and ambiguities. Only approval of the exact intent opens the second interface, which displays the assertion, checks, authorship, coverage warnings, and repair budget. If the client has no custom renderer, it receives these proposals through Harness’s native question UI. Approval in both stages is receipt-bound; failure at either stage leaves no active contract.

Code Verification and Evidence Management

The plugin runs code checks through Harness’s ctx.subprocess and ctx.sandbox; the model only supplies the profile ID. The plugin never uses the LLM as a result judge, starts no provider retry/fallback, and does not repeat business operations.

Evidence from trusted verifiers is conservatively invalidated. Any subsequent non-governor tool call (including nested code-mode scheduling), as well as any later different verification profile with workspace-write access, invalidates it. This prevents test-result validation from remaining attested after the agent or another verifier changes the code.

Installation and Usage

This is an unofficial community project under the MIT License. It is available through the community catalog or GitHub repository:

  • Community Catalog: https://www.skillhub.cn/plugins/chenjie1129/deepseek-harness-reliability-governor
  • GitHub Repository: https://github.com/chenjie1129/deepseek-harness-reliability-governor

The plugin version is 0.7.0 and requires Node.js version ^22.19.0 || >=24.0.0.

Notes

  1. Unofficial Project: This is a community project and is publicly recruiting public-beta testers. See FEEDBACK.md in the repository for the feedback protocol.
  2. Does Not Replace LLM Judgment: The plugin does not make the LLM deterministic; it only makes narrower commitments about agent guidance under contracts.
  3. Conservative Evidence Invalidation: Any subsequent non-governor tool call can invalidate trusted evidence and require validation to be rerun.

Summary

Through explicit contracts and evidence gating, the plugin provides infrastructure for deterministic completion and code verification in DSH agents. It is suitable for development scenarios that require strict evidence review and bounded testing of task outputs.