In the pluggable ecosystem of DeepSeek Harness, multi-model collaboration often struggles with inconsistent outputs, lack of objective review, or reliance on a single model’s judgment. The dsh-model-jury plugin aims to provide a structured cross-model peer-review mechanism that enables a group of models to reach consensus or identify disagreements through blind review, anonymous peer evaluation, and deterministic adjudication.
Core Features¶
The plugin implements the following workflow and capabilities:
- Structured cross-model peer review: It defines a complete state machine from question to adjudication.
- Blind independent response: In the first round, all models receive the same instruction and none can see another model’s response.
- Anonymous peer review: In the second round, models comment on anonymized peer responses without revealing identity.
- Revision stage: In the third round, models have an opportunity to revise their own answers based on the peer-review results.
- Deterministic adjudication: Code logic computes the voting results (majority, dissent, risks) and generates the final report.
- Provider-neutral: It is based on public services and does not implement private HTTP clients or proxies.
- No permanent chair: Adjudication is executed by code logic and does not depend on a particular model as chair.
Installation and Enablement¶
Use the official installation command to add the plugin to the DSH Web configuration file:
dsh plugin --profile web add dsh-model-jury
After installation, check whether the configuration takes effect, then start the Web interface:
dsh --profile web --dump-config
dsh web
Configuration and Model Settings¶
Model Jury uses the provider-neutral services provided by DSH. Different models require different configuration.
GPT (Codex)¶
The GPT seat uses DSH’s native Codex app-server provider and leverages the user’s native ChatGPT/Codex subscription authentication. OPENAI_API_KEY is not required. The plugin registers an isolated model-jury-codex route.
GLM and DeepSeek¶
These two providers require the user to manually configure credentials and model IDs. You need to override the model-jury configuration item in cordis.patch.yml:
- id: model-jury
config:
storageDir: .dsh-model-jury
maxQuestionChars: 12000
maxFieldChars: 6000
codex:
provider: model-jury-codex
timeoutMs: 180000
maxRetries: 1
glm:
provider: glm
model: '<your GLM model id>'
timeoutMs: 120000
maxTokens: 16384
maxRetries: 1
deepseek:
provider: deepseek-official
model: '<your DeepSeek model id>'
timeoutMs: 120000
maxTokens: 16384
maxRetries: 1
Credentials (such as GLM_API_KEY and DEEPSEEK_API_KEY) should be injected via DSH settings or the startup environment. They must never be included in the plugin package or patch files.
Usage Examples¶
Diagnose the Environment¶
Use the /jury doctor command to perform a minimal live-request test for all seats in the current environment, consuming quota and checking status:
/jury doctor
Regular Review¶
Start a review to answer a specific question:
/jury Should this inference runtime prioritize iOS support or ARM SIMD optimization?
Conflict Mode¶
Use the --style adversarial parameter to strengthen the peer-review instructions and force models to seek counterarguments:
/jury --style adversarial Should this inference runtime prioritize iOS support or ARM SIMD optimization?
The default style is balanced.
Workflow¶
The review process is divided into three rounds:
- Round 1 (blind review): All models receive the same question, instructions, and JSON Schema, and answer concurrently. The system randomly maps responses to the
P1,P2, andP3seats, but the models do not know their own position mapping. - Round 2 (peer review): Models receive their own Round 1 response and anonymized peer responses. The system strips provider, model names, and other identifying information; models only see anonymous content and must comment on the counterpart’s strengths, weaknesses, missing evidence, and actual disagreements.
- Round 3 (revision): Based on the peer-review results, models decide whether to revise, merge, maintain a dissenting position, or keep their original answer.
Finally, the deterministic aggregator computes the voting state based on the final states of the three parties (P1, P2, P3, hybrid, undecided) and generates the adjudication report.
Applicable Scenarios and Limitations¶
Applicable scenarios: Scenarios that require cross-validation of technical solutions by multiple models, checking for code logic vulnerabilities, or evaluating different decision paths.
Known limitations:
* Credential management: GLM and DeepSeek require the user to provide credentials and model IDs in the configuration.
* Sandbox limitations: The Codex provider cannot enforce a strict read-only sandbox through plugin configuration; this is a documented public limitation.
* No file/network operations: The review process is deliberation-only; it does not include file editing, dependency installation, submissions, or network mutation operations.
* DSH developer preview: DeepSeek Harness is in developer preview and may introduce breaking changes.
* Voting consensus limitation: The specific text on voting consensus limitations in the source document was truncated; consult the source code.
Summary¶
dsh-model-jury provides DeepSeek Harness with a decentralized model review protocol. By using code rather than humans as the judge, it ensures deterministic adjudication. For developers who need objective, multi-perspective technical assessments, this is a practical tool.
For plugin details and source code, visit: GitHub repository | Community catalog