Introduction¶
In agent development, the most expensive mistakes are often not faulty fixes, but false “closing conclusions.” A conclusion such as “it doesn’t work,” “there is no bug,” or “it cannot be reproduced,” if unsupported by evidence, leaves no trace and quietly closes the line of investigation. In contrast, an incorrect “it works” claim can be caught by later runs.
The dsh-plugin-verdict-guard plugin resolves this asymmetry mechanically: at the stop boundary of an agent turn, it checks whether the model’s impending closing conclusion contains verifiable evidence. If the conclusion is a verdict but lacks evidence, the turn is intercepted and the model is guided to supplement evidence.
Plugin Overview¶
dsh-plugin-verdict-guard is a native DeepSeek Harness (DSH) plugin that intercepts and steers turns that merely state conclusions without supporting evidence.
- Maintainer: sagetta1
- License: MIT
- Category: Model inference
- Core value: Ensures that the model must provide verifiable evidence (such as tool execution results) before closing a direction, preventing silent misjudgments.
Core Features¶
- Intercept unsupported conclusions: When the model tries to end a turn and issues a verdict-style conclusion (such as “it doesn’t work”) without providing any inspectable content, the turn is intercepted.
- Verify tool results: The plugin checks not only the model’s textual output, but also the actual tool call results in the session log (such as
bash,grep, andread). If the model asserts that something is true but no verification tool was actually run, the conclusion is treated as invalid. - Distinguish paths from evidence: The plugin distinguishes “file paths” (which may only be recalled) from “command transcripts” (such as fenced blocks, test results, and HTTP status codes). The former is not sufficient to trigger strict interception, while the latter can.
- Limit intervention frequency: By default, the plugin intercepts only once per turn. This is a “speed bump” rather than a deadlock mechanism: it prevents infinite loops while also limiting the total number of interventions in a single session.
Installation and Configuration¶
The plugin is installed through DSH’s plugin manager.
Installation¶
dsh plugin add dsh-plugin-verdict-guard
Verify Installation¶
After installation, check that the plugin has been correctly merged into the configuration tree:
dsh --profile headless --dump-config | grep -A 3 verdict-guard
Configuration¶
The plugin supports configuration via cordis.patch.yml. The following shows the default configuration and examples of common modifications:
- id: verdict-guard
config:
locale: en
requireToolEvidence: false
maxInterventionsPerTurn: 2
Configuration Options¶
| Option | Default | Description |
|---|---|---|
locale |
both |
Language option for the verdict vocabulary; supports en, ru, or both |
requireToolEvidence |
true |
Whether a verdict requires verifying tool results (such as bash/read/grep), rather than any tool result |
maxInterventionsPerTurn |
1 |
Maximum number of interceptions allowed in a single turn |
maxInterventionsPerSession |
6 |
Maximum number of interceptions allowed in a single session |
verbose |
false |
Whether to log all pass/intercept decisions at debug level |
How It Works¶
The plugin listens to DSH’s agent/turn-stopping event. At the turn stop boundary, it performs the following steps:
1. Reads the closing text the model is about to output.
2. Reads the tool calls and tool results that actually occurred during the turn (session log).
3. Determines whether the text contains verdict vocabulary and whether corresponding evidence exists.
4. If no evidence exists, calls agent.steer(reason) to force the model to re-examine the input and run the next step until the turn ends or evidence appears.
Important note: DSH currently has a TODO(stop-loop-guard) in its loop control, so the listener must limit the number of interventions itself (via maxInterventionsPerTurn); otherwise, it may cause an infinite loop.
Caveats and Compatibility¶
- No truth verification: The plugin only checks whether inspectable evidence exists; it does not verify whether the conclusion itself is true (it does not act as an Oracle).
- Tool dependency: The plugin requires verification tools (such as bash, read, or grep) to generate evidence. Simply writing files is generally not considered verification.
- Version compatibility: The plugin is built and tested against
0.1.0-rc.8.- ⚠️ The
latesttags for DSH’s own subpackages are outdated. Use the@nexttag or explicitly specify the^0.1.0-rc.8range to avoid mixing incompatible versions.
- ⚠️ The
- Permissions and provenance: Review the source code and license before installation.
Summary¶
dsh-plugin-verdict-guard effectively curbs silent misjudgments by introducing an evidence-checking mechanism at the turn stop boundary. It does not replace human review; instead, it acts as an automated “speed bump,” requiring the model to provide a traceable verification process before closing a problem direction.
Catalog page: dsh-plugin-verdict-guard - SkillHub
Source repository: GitHub - sagetta1/dsh-verdict-guard