The core philosophy of DeepSeek Harness (DSH) is “everything is a plugin.” When building agents or orchestrating model calls, developers often need fine-grained control over sampling parameters for specific models or provider routes. The dsh-llm-sampling plugin provides this capability by enforcing sampling policies through the agent/request waterfall.
Plugin Positioning¶
This is a DeepSeek Harness plugin package maintained by kuma-loong. It manages deployment-level sampling policies. The DSH core handles provider-neutral fields and persistent request headers, while adapters handle wire translation; this plugin only manages the injection of sampling policies.
Installation¶
When installing from GitHub, pin a reviewed commit to ensure environment consistency.
dsh plugin --profile web add github:kuma-loong/dsh-llm-sampling#<commit>
Installing from Git runs the prepare script inside the package. If you use pnpm 10 or later, you must explicitly allow builds in the profile’s pnpm-workspace.yaml:
allowBuilds:
dsh-llm-sampling@https://codeload.github.com/kuma-loong/dsh-llm-sampling/tar.gz/<commit>: true
Copy the exact key printed by pnpm, then rerun the installation command.
Configuration¶
Add an llm-sampling configuration section to $DSH_HOME/settings.yaml. Supported sampling fields include temperature, topP, topK, minP, presencePenalty, and repetitionPenalty.
llm-sampling:
providers:
sparse-vllm:
models:
Qwen3.8-27B:
default:
temperature: 1
topP: 0.95
topK: 20
minP: 0
presencePenalty: 0
repetitionPenalty: 1
off:
temperature: 0.7
topP: 0.8
presencePenalty: 1.5
The Profiles here are policies, not caller-side defaults: configured values override early proposals in the agent/request waterfall. Subsequent request policies can also intentionally override them through the normal waterfall order.
How It Works¶
The plugin activates after a route is named in llm-sampling.providers. For a named route, the sampling values from the configured model’s default Profile replace all sampling values. When reasoningEffort is explicitly set to off, the off Profile fully overrides the default Profile.
The plugin does not add extra text or tool schemas to the model prompt. It only injects sampling values into requests sent to the model provider and records these effective values in the session’s request/header. This affects the generation distribution and may change output length. Because the prompt prefix remains unchanged, KV caches are not invalidated by sampling policy changes. However, providers may include sampling controls in the request cache identity, so policy changes can affect provider-side cache reuse.
Limitations and Requirements¶
- Dependency Requirements: The plugin requires a build of DeepSeek Harness in which
LlmCallConfigandGenerateOptionsincludetopP,topK,minP,presencePenalty, andrepetitionPenalty. Until the corresponding npm release meets these requirements, only a source build that includes these changes can be used. - Adapter Mapping: The adapter must be able to map the configured fields. For example, the
@deepseek-ai/dsh-llm-pi-aiadapter supports extended fields for OpenAI Chat Completions, but other protocols may reject these fields. - Dormant State: Before any route is named in
llm-sampling.providers, the plugin is dormant and has no effect on requests. - Unconfigured Routes: Requests on routes that are not configured remain unchanged and are not affected by the plugin.
This plugin provides fine-grained parameter control for model inference. For more details, see the plugin catalog or the source repository.