In the pluggable ecosystem of DeepSeek Harness, multi-model collaboration often struggles with inconsistent outputs, lack of objective review, or reliance on a single model’s judgment. The dsh-model-jury plugin aims to provide a structured cross-model peer-review mechanism that enables a group of models to reach consensus or identify disagreements through blind review, anonymous peer evaluation, and deterministic adjudication.

Core Features

The plugin implements the following workflow and capabilities:

  • Structured cross-model peer review: It defines a complete state machine from question to adjudication.
  • Blind independent response: In the first round, all models receive the same instruction and none can see another model’s response.
  • Anonymous peer review: In the second round, models comment on anonymized peer responses without revealing identity.
  • Revision stage: In the third round, models have an opportunity to revise their own answers based on the peer-review results.
  • Deterministic adjudication: Code logic computes the voting results (majority, dissent, risks) and generates the final report.
  • Provider-neutral: It is based on public services and does not implement private HTTP clients or proxies.
  • No permanent chair: Adjudication is executed by code logic and does not depend on a particular model as chair.

Installation and Enablement

Use the official installation command to add the plugin to the DSH Web configuration file:

dsh plugin --profile web add dsh-model-jury

After installation, check whether the configuration takes effect, then start the Web interface:

dsh --profile web --dump-config
dsh web

Configuration and Model Settings

Model Jury uses the provider-neutral services provided by DSH. Different models require different configuration.

GPT (Codex)

The GPT seat uses DSH’s native Codex app-server provider and leverages the user’s native ChatGPT/Codex subscription authentication. OPENAI_API_KEY is not required. The plugin registers an isolated model-jury-codex route.

GLM and DeepSeek

These two providers require the user to manually configure credentials and model IDs. You need to override the model-jury configuration item in cordis.patch.yml:

- id: model-jury
  config:
    storageDir: .dsh-model-jury
    maxQuestionChars: 12000
    maxFieldChars: 6000
    codex:
      provider: model-jury-codex
      timeoutMs: 180000
      maxRetries: 1
    glm:
      provider: glm
      model: '<your GLM model id>'
      timeoutMs: 120000
      maxTokens: 16384
      maxRetries: 1
    deepseek:
      provider: deepseek-official
      model: '<your DeepSeek model id>'
      timeoutMs: 120000
      maxTokens: 16384
      maxRetries: 1

Credentials (such as GLM_API_KEY and DEEPSEEK_API_KEY) should be injected via DSH settings or the startup environment. They must never be included in the plugin package or patch files.

Usage Examples

Diagnose the Environment

Use the /jury doctor command to perform a minimal live-request test for all seats in the current environment, consuming quota and checking status:

/jury doctor

Regular Review

Start a review to answer a specific question:

/jury Should this inference runtime prioritize iOS support or ARM SIMD optimization?

Conflict Mode

Use the --style adversarial parameter to strengthen the peer-review instructions and force models to seek counterarguments:

/jury --style adversarial Should this inference runtime prioritize iOS support or ARM SIMD optimization?

The default style is balanced.

Workflow

The review process is divided into three rounds:

  1. Round 1 (blind review): All models receive the same question, instructions, and JSON Schema, and answer concurrently. The system randomly maps responses to the P1, P2, and P3 seats, but the models do not know their own position mapping.
  2. Round 2 (peer review): Models receive their own Round 1 response and anonymized peer responses. The system strips provider, model names, and other identifying information; models only see anonymous content and must comment on the counterpart’s strengths, weaknesses, missing evidence, and actual disagreements.
  3. Round 3 (revision): Based on the peer-review results, models decide whether to revise, merge, maintain a dissenting position, or keep their original answer.

Finally, the deterministic aggregator computes the voting state based on the final states of the three parties (P1, P2, P3, hybrid, undecided) and generates the adjudication report.

Applicable Scenarios and Limitations

Applicable scenarios: Scenarios that require cross-validation of technical solutions by multiple models, checking for code logic vulnerabilities, or evaluating different decision paths.

Known limitations:
* Credential management: GLM and DeepSeek require the user to provide credentials and model IDs in the configuration.
* Sandbox limitations: The Codex provider cannot enforce a strict read-only sandbox through plugin configuration; this is a documented public limitation.
* No file/network operations: The review process is deliberation-only; it does not include file editing, dependency installation, submissions, or network mutation operations.
* DSH developer preview: DeepSeek Harness is in developer preview and may introduce breaking changes.
* Voting consensus limitation: The specific text on voting consensus limitations in the source document was truncated; consult the source code.

Summary

dsh-model-jury provides DeepSeek Harness with a decentralized model review protocol. By using code rather than humans as the judge, it ensures deterministic adjudication. For developers who need objective, multi-perspective technical assessments, this is a practical tool.

For plugin details and source code, visit: GitHub repository | Community catalog