Introduction¶
When deploying large models locally, developers often face cumbersome model switching and disorganized process management. Connecting DeepSeek Harness (DSH) directly to llama-server is feasible, but when dynamically switching models, it usually requires manually stopping the service, modifying the configuration, and restarting it. This not only adds operational overhead but can also lead to interrupted requests or improperly released VRAM. The dsh-llama-model-manager plugin addresses this pain point by using a stable OpenAI-compatible gateway to automatically manage the loading and unloading of local models.
Plugin Overview¶
This plugin is a DeepSeek Harness (DSH) component specifically designed to run llama-server.exe in a Windows environment and expose local GGUF models to DSH through an OpenAI-compatible interface. It takes over the entire model lifecycle—including startup, shutdown, switching, and fault recovery—so DSH only needs to connect to a fixed URL without needing to concern itself with underlying model changes.
Core Features¶
- Lifecycle Management: Automatically handles model loading and unloading. When a model switch is required, the plugin first stops the old model and then loads the new one.
- Stable Gateway: Provides a fixed OpenAI-compatible gateway address (default port 8080), so DSH does not need frequent configuration changes.
- Graceful Shutdown and VRAM Release: Uses a
Ctrl+Csignal rather than forcefully terminating the process, ensuring thatllama.cppcan clean up resources properly and avoid VRAM leaks. - Hidden Console: Hides the console window when the process starts, avoiding flicker and visual interference.
- Orphan Process Handling: Records process state and provides safety mechanisms to prevent resource leaks left behind after a DSH crash.
- Visual Settings: Provides a settings page that displays status, logs, and the model list in real time, and supports operations directly from the interface.
Installation and Enabling¶
Ensure the environment meets the following requirements:
* Operating System: Windows 10/11.
* Runtime Environment: Node.js >= 20.10.
* Dependencies: DSH Web UI, and llama-server.exe (from llama.cpp).
The installation command is as follows:
dsh plugin --profile web add github:DoctorxPriestess/dsh-llama-model-manager
After installation, DSH must be restarted to load the plugin.
Configuration and Usage¶
In the DSH settings page, enter the path to llama-server.exe and the model information. Then configure the Provider in DSH’s settings.yaml to point to the plugin’s gateway address:
llm-pi-ai:
providers:
llamacpp:
displayName: llama.cpp local
api: openai-completions
apiKeyEnv: LLAMACPP_API_KEY
baseURL: http://127.0.0.1:8080/v1
models:
- id: qwen38-iq3s
name: Qwen3.8-27B IQ3_S
contextWindow: 131072
Port Description:
* Gateway port: 8080 (DSH connection address).
* Internal port: 18080 (llama-server listening address).
Stop Method:
The plugin supports controlling stop behavior by setting stopMethod:
* auto (default): Sends Ctrl+C first; force-terminates after timeout.
* ctrl-c: Only sends Ctrl+C and does not force-terminate.
* taskkill: Skips Ctrl+C and force-terminates directly.
Notes¶
- Platform Limitation: The plugin only supports Windows 10/11.
- Permissions and Security: The plugin runs with the permissions of the DSH process; it is recommended to review the source code before installation.
- Configuration Isolation: The plugin itself does not read or write DSH’s
settings.yaml, but thebaseURLmust be correctly filled in when configuring the Provider.
Summary¶
By automating lifecycle management, dsh-llama-model-manager resolves resource leaks and cumbersome operations during local model switching, providing developers with a stable and easy-to-manage local inference solution.