Introduction¶
When running multiple inference services locally (such as llama.cpp, SGLang, or vLLM), configuration parameters are often scattered across terminal history and notebooks: which set of -c / -np parameters should a given model use? Is the VRAM sufficient? Which framework can actually start on a specific GPU?
dsh-model-manager is a plugin for DeepSeek Harness (DSH), designed to provide a unified “control console” for local LLM inference services. With it, you can manage service registration, parameter profiles, GPU status, and benchmarking from a single panel.
Core Features¶
This plugin mainly addresses the challenges of chaotic local multi-service management, difficult VRAM validation, and hard-to-trace parameter configuration.
-
Service Registry
- Register local inference services and record port, framework, model, and GPU information.
- Provide real-time health checks and support one-click stop.
-
Parameter Profile Management
- Manage named parameter versions for llama.cpp / SGLang / vLLM.
- Parameter tables are fixed-sorted into 9 logical groups (model → context → KV → speculative decoding → sampling → performance → parallelism → server → other).
- Different quantization versions of the same model × same framework are displayed as a union, making it easy to compare differences.
-
GPU Detection
- Automatically enumerate GPUs, preferring
libcudaand falling back tonvidia-smiif missing. - Card indices correspond to
CUDA_VISIBLE_DEVICESvalues.
- Automatically enumerate GPUs, preferring
-
VRAM Validation
- When saving a Profile, estimate KV usage based on
-c/-npparameters. - If the estimated usage exceeds the target card capacity, a warning is issued.
- When saving a Profile, estimate KV usage based on
-
One-Click Benchmark
- Send a fixed prompt to a running service (non-streaming, 256 tokens) and record tok/s.
- Support full-context hot benchmarking (verify available context ≥ target, and record TTFB / prefill / decode).
-
Safety Guardrails
- Never execute
pkill. - Stopping external services requires explicitly selecting
forceand performing two clicks to confirm in the panel. - Stopping port 11437 (DSH’s own inference port) will prompt “will interrupt the current session”.
- Never execute
Installation and Enabling¶
This plugin is built on DSH’s plugin ecosystem.
- Enter the DSH web profile directory:
cd ~/.dsh/profiles/web
- Install the plugin using pnpm:
pnpm add github:Ansonfishing/dsh-model-manager
-
Modify
package.jsonand add"dsh-model-manager"to thedsh.profile.bundlesarray. -
Restart
dsh. A “Model Management” tab will appear in the session view.
Typical Usage¶
Configure the Local GPU Table¶
To make the panel display actual GPU names and VRAM capacity instead of the default “GPU 0” / “GPU 1”, you can create a configuration file:
touch ~/.dsh/model-manager/builtin-gpus.local.json
Then write the following content into the file:
{
"0": { "name": "RTX 4090", "memGb": 24 },
"1": { "name": "RTX 6000 Ada", "memGb": 48 }
}
If this file does not exist, the panel will display default names, and the VRAM validation feature will be automatically skipped.
Scenarios and Notes¶
- Dependencies: DSH (with web interface) + pnpm is required. GPU detection requires python3 (falling back to nvidia-smi if missing).
- Permissions and Security: The plugin runs with the privileges of the current DSH process. Stopping external services includes a strict secondary confirmation mechanism to prevent accidental operations.
- License: MIT License.
Summary¶
dsh-model-manager is a practical tool for local LLM developers. By providing a unified registry and parameter management system, it solves the configuration chaos caused by running multiple services in parallel. For developers who need fine-grained VRAM control, benchmarking, and safe service start/stop operations, this plugin provides the necessary support.