Foreword

DeepSeek Harness (DSH) allows each chat session to independently select a model, and background assistants can also request other models. On machines that can only load a single local model, these “out-of-bounds” requests can evict the loaded model and trigger the llama.cpp router to reload. This is not only time-consuming (several minutes), but also causes loss of the prompt cache. In addition, subagents inherit the model from when the session was created, rather than the model currently in use, which can lead to routing errors. The dsh-model-pin plugin implements an allowlist at DSH’s parsing layer (agent/request pipeline) and intercepts requests that do not meet the requirements.

Plugin Overview

dsh-model-pin is a DeepSeek Harness plugin maintained by d3vmeh. It addresses the “model roulette” problem: model requests per provider must remain within an allowed set, mismatched requests are redirected or rejected, and a warning is issued when the llama.cpp router reloads. This plugin depends on @deepseek-ai/schemastery.

Core Features

  • Constrains per-provider model requests to an allowed set.
  • Redirects or rejects mismatched requests.
  • Issues a warning when the llama.cpp router reloads.
  • Implements an allowlist at the DeepSeek Harness parsing layer.
  • Prevents subagents from inheriting an incorrect model.
  • Prevents prompt cache loss (by avoiding model reloads).

Installation and Enablement

Install using the official command:

dsh plugin --profile web add dsh-model-pin

After installation, it needs to be enabled in the configuration file. The configuration file is usually located at ~/.dsh/profiles/web/cordis.patch.yml.

Typical Usage

Add a model-pin plugin section in the configuration file. An example configuration is as follows:

- id: model-pin
  config:
    providers:
      llamacpp:
        allow: [qwen3.8-q4-long, qwen3.8-fast]
        fallback: qwen3.8-q4-long
        action: redirect
        warnOnSwitch: true
  • allow: Specifies the list of allowed models. If only one entry is retained, single-model mode is enabled.
  • fallback: The default model, usually the first entry in the allow list.
  • action: Defines how mismatched requests are handled. redirect (default) or reject.
  • warnOnSwitch: Specifies whether to issue a warning when a request may cause a model switch.

Use Cases and Notes

  • Resource-Constrained Environments: When a machine can only retain one local model, this plugin can avoid model reloads and cache loss caused by unintended requests.
  • Strict Routing Control: Ensures that all subagents and background tasks consistently use a specific model.
  • Observational Nature: Switch warnings are intended only for observation and cannot prevent reloads. To completely prohibit reloading, set allow to a single-model list.
  • Web Interface Differences: The web model selector is not aware of the pin and may display an incorrect model. Check terminal logs or /logs for actual runtime behavior.
  • Inheritance Mechanism: Subagents inherit the model from when the session was created, rather than the model currently used in the session. This plugin intercepts at session boundaries.
  • Cross-Provider Support: Version 1 does not support cross-provider redirection.
  • Precedence: The plugin is registered when the profile is loaded and has final decision authority, but other root plugins that load earlier may override it.
  • License: MIT.

Summary

dsh-model-pin resolves performance overhead and cache loss caused by frequent model switching by enforcing an allowlist at the DSH parsing layer. For scenarios requiring stable model routing or resource constraints, it is a practical constraint tool. See the GitHub repository for more details.