Introduction¶
DeepSeek Harness (DSH) provides the underlying capabilities for building agents. During development and debugging, obtaining cross-provider model output, validating routing logic, or performing audits are common requirements. dsh-consult is a local auxiliary plugin that provides DSH with a cross-provider “second opinion” service while also exposing an OpenAI-compatible local API endpoint, allowing external Harnesses (such as ZCode) to call lightweight advisor functionality.
Unlike the passive dsh-advisor, it is an active consulting tool.
Core Features¶
This plugin mainly provides the following capabilities:
- DSH_MODEL direct model selection: Routes traffic via environment variables or configuration to force a specific model for consultation, without being overridden by the default model in the user-level
settings.yaml. - OpenAI-compatible endpoint: Provides the standard
POST /v1/chat/completionsinterface at127.0.0.1:3080, supporting SSE streaming and non-streaming responses, as well as structured output. - consult tool: Registered for all agents and used to actively request cross-provider opinions during a session.
- scout investigation mode: Combines official models with real-time Web Search to resolve difficult problems that neither humans nor models can handle easily.
- Structured output: Requires JSON output in the form
{recommendation, reasons, confidence, risks, alternatives}by default, withjson_schemaconstraints supported. - Upstream tracing and auditing: Response headers identify the upstream model, and audit logs are persisted to
audit.jsonl, including call latency and routing information.
Installation and Enablement¶
Installation requires manually cloning the repository and modifying the profile configuration files.
- Clone the repository:
git clone https://github.com/1339190177/dsh-consult ~/.dsh/local-plugins/dsh-advisor
- Modify Profile Configuration:
Enter the DSH web and headless profile directories, editpackage.json, and add the dependency and bundle reference:
// ~/.dsh/profiles/web/package.json
{
"dependencies": {
"dsh-advisor": "file:../local-plugins/dsh-advisor"
},
"dsh": {
"profile.bundles": ["dsh-advisor"]
}
}
// ~/.dsh/profiles/headless/package.json (same structure)
- Install Dependencies:
Run the following separately in the web and headless directories:
cd ~/.dsh/profiles/web && pnpm install
cd ~/.dsh/profiles/headless && pnpm install
The web profile takes effect after restarting `dsh web`, while the headless profile takes effect immediately on each run.
Configuration Notes¶
After installation, you need to enable the HTTP endpoint and configure an authentication token.
Bundle-layer configuration (cordis.patch.yml):
- id: dsh-advisor
name: 'dsh-advisor'
config:
httpEnabled: true # If disabled, only the consult tool is retained
User-layer configuration (~/.dsh/settings.yaml):
advisor:
token: dsha_... # Automatically generated on first launch, used for HTTP authentication
route: # Optional: fix the default routing for consult
provider: deepseek
model: deepseek-v4-flash
Usage¶
The plugin supports multiple invocation methods.
1. CLI Mode (Headless)¶
Specify the model via the DSH_MODEL environment variable and invoke it directly from the command line.
# Directly select the model and consult
DSH_MODEL=deepseek/deepseek-v4-flash dsh --profile headless "question"
# Verify that the configuration and selection are consistent
DSH_MODEL=deepseek/deepseek-v4-flash dsh --profile headless --dump-config
2. HTTP API Mode¶
Call the local endpoint using the standard OpenAI format.
curl -X POST http://127.0.0.1:3080/v1/chat/completions \
-H 'content-type: application/json' \
-H 'authorization: Bearer dsha_...' \
-d '{"model":"deepseek/deepseek-v4-flash","messages":[{"role":"user","content":"…"}]}'
Streaming ("stream":true) and structured output (response_format: json_schema) are supported.
3. Scout Investigation Mode (Multimodal)¶
The script ask_advisor.sh supports image input and Web Search.
bash ask_advisor.sh "question" "context" image1.png image2.png
Notes¶
- Reasoning Effort: By default, no parameter is passed, so the provider’s default setting is followed (the official flash model defaults to high and adaptive).
offis the only option that can force reasoning off. When using resellers such as XQAPI, this parameter may be ineffective; it is recommended to connect directly to official models. - Unimplemented Features: F7 (quota limiting) and F8 (concurrent quorum) are currently unimplemented.
- Parameter Limitations: The CLI
--modelflag itself requires modifying the DSH source code (an issue should be filed). This plugin achieves more stable model selection by using theDSH_MODELenvironment variable together withagent/requestrouting. - Security Boundary: The HTTP endpoint only accepts loopback connections and does not enable CORS. In production environments, it is recommended to configure API-layer Origin same-origin validation.
- Audit Backfill: The
adoptedfield returned by the consult tool currently relies on the main model marking “adopted” or “not adopted” in the response text; automatic backfill is not yet implemented.