Introduction¶
The core design philosophy of DeepSeek Harness is “everything is a plugin.” When developers need to integrate models or services that are not natively supported, modifying core code is often required. The dsh-llm-openai-compatible plugin solves this problem. As a “universal plug,” it uses the OpenAI-compatible protocol to allow Harness to connect directly to any local or remote OpenAI-compatible endpoint.
Plugin Overview¶
Name: dsh-llm-openai-compatible
Maintainer: cqnxnzg
Category: admin-security
License: MIT
The core value of this plugin lies in decoupling. After installation, by configuring only the endpoint address and model catalog, Harness can communicate with the compatibility layers of vLLM, LM Studio, llama.cpp, and Ollama, or with remote gateways such as OpenRouter, Together, and Moonshot. It does not require installing Ollama, nor does it require modifying the core code of DeepSeek Harness.
Installation and Enablement¶
Before installing, ensure that you have installed DeepSeek Harness 0.1.0-rc.6 or a newer version.
dsh plugin --profile web add github:cqnxnzg/dsh-llm-openai-compatible
After installation, the plugin creates a configuration entry named llm-openai-compatible under Settings -> Plugins.
Core Features¶
- Any OpenAI-compatible endpoint: Supports local servers or remote gateways.
- API key is optional: Local endpoints support anonymous requests; remote gateways usually require authentication.
- Model catalog declaration: Declares the models actually provided by the endpoint and their attributes through the
modelsarray. - Web configuration interface: Allows visual configuration changes through Settings -> Plugins ->
llm-openai-compatible. Changes take effect immediately after saving. - Model discovery: Reads the endpoint’s
GET <base>/v1/modelsinterface throughdiscoverModels()to automatically obtain the model list. - Installation without allowlist: The repository directly commits the build artifacts
lib/, so it can be installed directly without the pnpm build-script allowlist.
Configuration Guide¶
Go to Settings → Plugins → llm-openai-compatible. All configuration fields are optional.
Main Configuration Fields¶
| Field | Default Value | Description |
|---|---|---|
apiKeyEnv |
OPENAI_API_KEY |
Credential reference (environment variable name); if not configured or empty, local endpoints can request anonymously |
baseURL |
http://127.0.0.1:8000/v1 |
OpenAI-compatible endpoint; the plugin normalizes a trailing /v1 automatically |
models |
4 sample models | Model catalog actually served by the endpoint; unlisted models cause UNKNOWN_MODEL requests |
maxTokens |
— | Global default output limit; used as a fallback when a model row does not declare one |
defaultContextWindow |
131072 |
Context capacity when the model does not declare contextWindow |
streamIdleTimeoutMs |
300000 |
Idle timeout for streaming reads |
retryPolicy |
normal default | Retry policy configuration |
Model Catalog Fields¶
Each item in the models array contains the following fields:
| Field | Description |
|---|---|
id |
Model id accepted by the endpoint (must match the id actually served by the endpoint, otherwise an error is returned) |
name |
Display name in the selector; if omitted, id is used |
description |
Additional description in the selector (optional) |
contextWindow |
Context capacity for this model (tokens) |
maxTokens |
Model-specific output limit, with higher priority than the global maxTokens |
vision |
true indicates that image input is accepted |
thinking |
true indicates native thinking support; must be used together with thinking levels in the selector |
defaultEffort |
Default thinking level in the chat selector |
tools |
Legacy capability flag; ignored at runtime |
Retry Policy Configuration¶
The default policy is normal mode: up to 2 retries, with backoff retries for error codes such as EMPTY_RESPONSE, RATE_LIMIT, SERVER, TIMEOUT, and TRANSPORT.
retryPolicy:
mode: normal # normal | always
maxRetries: 3 # normal 模式下的最大重试次数
retryableCodes: [RATE_LIMIT, SERVER, TIMEOUT, TRANSPORT]
backoff:
initialDelayMs: 500
maxDelayMs: 10000
jitterRatio: 0.1
Common Endpoint Configuration Examples¶
| Service | baseURL Example | Notes |
|---|---|---|
| vLLM | http://127.0.0.1:8000/v1 |
The name specified by --served-model-name is the request id |
| LM Studio | http://127.0.0.1:1234/v1 |
— |
| llama.cpp server | http://127.0.0.1:8080/v1 |
— |
| Ollama (compat layer) | http://127.0.0.1:11434/v1 |
Must use the compatibility layer; model ids often include a tag, such as qwen2.5:7b |
| OpenRouter | https://openrouter.ai/api/v1 |
Gateway, must configure apiKeyEnv |
| Together | https://api.together.xyz/v1 |
Gateway, must configure apiKeyEnv |
Connecting to the Official DeepSeek API¶
The official DeepSeek endpoint is itself OpenAI-compatible.
llm-openai-compatible:
apiKeyEnv: DEEPSEEK_API_KEY # 环境变量名
baseURL: https://api.deepseek.com # 自动路由到 /v1
models:
- id: deepseek-chat # 通用对话
name: DeepSeek Chat
contextWindow: 65536
maxTokens: 8192
- id: deepseek-reasoner # 推理模型
name: DeepSeek Reasoner
contextWindow: 65536
maxTokens: 8192
thinking: true # 必须声明支持思考
Troubleshooting¶
UNKNOWN_MODEL: The requested model id is not in the models catalog. Use curl http://127.0.0.1:8000/v1/models to obtain the actual model id returned by the endpoint, and update the configuration.
401 Unauthorized: Remote gateways usually require authentication. Check whether apiKeyEnv is configured correctly. For local endpoints that do not require a key, this field can be left empty.
404 / Cannot Connect: baseURL should include only the domain and port; do not include the full path (the plugin automatically appends /v1). Ensure the service is running and the port is correct.
/v1 Written Twice: Do not write .../v1/v1 in baseURL.
Development and Build¶
The plugin is developed in TypeScript. The committed code includes compiled artifacts lib/, so it can be installed directly without pnpm’s prepare script.
pnpm install
pnpm run build # tsc + tsdown