DSH (DeepSeek Harness) adopts an “everything is a plugin” architecture. In model inference scenarios, how to improve output quality without increasing compute cost is a core concern for developers. The aispin-dev/llm-as-a-verifier-dsh plugin encapsulates the test-time selection method of LLM-as-a-Verifier (arXiv:2607.05391) as a DSH-native plugin. It provides test-time scaling for DeepSeek V4 Flash through a Best-of-N dialogue mode: on the Terminal-Bench 2.1 dataset, Bo5 self-verification reaches 88% accuracy, reproducing frontier model performance at roughly 1/11 of the cost.

Core Principle

The core of this method is to solve the problem that “cheap models can generate good answers, but cannot recognize which one is good.” The plugin automates this pipeline:

  1. Sample candidates: Sample N slightly different answers from the configured model (e.g., DeepSeek V4 Flash).
  2. Fine-grained scoring: Use the same model to score candidates pairwise. The scoring is not simple letter sampling, but the expected value (Σ p(token)·φ(letter)) based on the log-probability distribution of grade tokens. This is the model’s complete belief, rather than a single random roll.
  3. Remove bias: Each candidate pair is scored bidirectionally to cancel out the verifier’s positional bias.
  4. Select the best output: In Best-of-N mode, the system automatically selects the best candidate to present to the user.

Installation and Enabling

The installation process depends on the npm package manager; resolving dependencies via the profile context is recommended.

dsh plugin --profile <your-profile> add @aispin/plugin-verifier

The plugin supports zero-configuration startup and inherits the provider state in the current dsh configuration (including credentials and settings). As long as DeepSeek has been configured on the Models page, the plugin can be used directly.

Features

The plugin provides three interaction modes for different scenarios:

  • Tool mode: the verify tool. When the agent in the conversation explicitly instructs “use the verify tool to compare A/B/C”, the plugin intervenes to perform verification.
  • Service mode: ctx.verifier.verify({ task, candidates }). For use by other plugins or orchestration scripts to control verification logic at the code level.
  • Automatic mode: Best-of-N dialogue mode. In a session where this mode is selected, the system automatically performs N-way sampling and verification in the background, and outputs only the winner to the user.

Implementation Details

The plugin natively implements the paper’s algorithm in DSH and includes the following features:

  • Fine-grained reward: Compute expected scores based on the top-20 logprob distribution, using a grouped grading scale (A=20…T=1). The scoring temperature is set to 1.0 to preserve a natural distribution.
  • PPT probability pivot tournament: Through a random Hamiltonian cycle and top-k pivot selection, only non-pivot-to-pivot pairs are completed, significantly reducing the number of scoring calls (e.g., Bo-5 requires only 11 scores).
  • Prefix-cache layout: Optimize the prompt structure by placing the role, grading scale, task, and candidates in a shared prefix to reduce token consumption.
  • Capability-adaptive scoring: Automatically detect whether an endpoint supports logprobs. Endpoints that support logprobs (such as the official endpoint) use expected scores; endpoints that do not (such as some gateways) automatically fall back to sampled pairwise scoring.
  • Fail-open discipline: If any sampling fails or times out, the system falls back to a normal response with an explanation, ensuring the conversation flow is not interrupted.

Configuration and Consumption

In the Web settings panel, you can control verification intensity via tiers:

Tier Model calls Token consumption Latency
Off 1 1× 1×
Fast · Bo-3 ~9 2–3× ~7–15s
Precise · Bo-5 ~16 3–5× ~12–30s
Custom 2–8 candidates Linear Linear

Sampling degradation chain: Each candidate sample has an independent budget. If two candidates time out during Bo-5 sampling, the system automatically downgrades to Bo-3 and continues selecting among the survivors. The actual usage is shown in the response footer (e.g., 5 candidates sampled, 2 incomplete · 1 of 3 selected).

Notes

  • Permissions: The plugin runs with the permissions of the current DSH process; check the source code and license before installation.
  • License: The implementation and methodology are both under the MIT license.
  • Unofficial: The plugin is not affiliated with the paper’s authors or DeepSeek.

The plugin provides DSH users with a low-cost, highly interpretable test-time scaling solution, suitable for production environments that require high reasoning quality while controlling costs.