Preface¶
In the plugin ecosystem of DeepSeek Harness, local model inference is a key scenario for reducing latency and protecting privacy. The dsh-local-llm plugin is specifically designed to manage local GGUF models and provide an OpenAI-compatible inference interface through llama-server, eliminating the need to configure Ollama.
Introduction¶
This is a DeepSeek Harness local LLM integration plugin maintained by user wertyBSd. It allows users to directly download, manage, and run models in GGUF format within the Harness environment, integrating the local LLM service as a Provider in Harness.
Core Features¶
- Model Download and Management: Supports downloading GGUF models from Hugging Face or direct URLs, searching the built-in model catalog, displaying the list of downloaded models and their file sizes, and deleting downloaded models.
- Download Optimization: Sends download progress through SSE streaming, supports concurrent downloads with deduplication, and automatically cleans up incomplete downloads and handles HTTP redirects.
- Security and Fault Tolerance: Prevents path traversal using file names and returns a clear error when the server is not ready.
- Smart Adaptation: Automatically selects an appropriate context size for the model and detects whether TinyLlama can accommodate Harness prompts before sending requests.
- Status Indication: Provides a server running status indicator in the sidebar footer.
Installation and Configuration¶
The system requires Node.js 18 or higher, and DeepSeek Harness must have the llm and webServer services enabled.
The installation command is as follows:
dsh plugin --profile web add https://github.com/wertyBSd/dsh-local-llm
The plugin includes a pre-built dist/ directory. A configuration example (which can be modified in dsh --dump-config) is as follows:
{
"model": "mistral-7b-instruct-v0.3-Q4_K_M.gguf",
"modelPath": "",
"runtimeUrl": "http://127.0.0.1:8080",
"contextSize": 8192,
"autoContextSize": true,
"port": 8080,
"downloadDir": "./models"
}
Parameter descriptions:
* model: Built-in model name or direct URL of a .gguf file.
* runtimeUrl: URL of the local inference service (default is http://127.0.0.1:8080).
* contextSize: Minimum context size passed to llama-server (default 8192).
* autoContextSize: Automatically selects the context size required by the model (enabled by default).
The built-in models include: mistral-7b-instruct-v0.3-Q4_K_M.gguf, llama-3-8b-instruct-q4_K_M.gguf, deepseek-coder-6.7b-instruct-q4_K_M.gguf, and qwen-2.5-7b-instruct-q4_K_M.gguf.
Usage Guide¶
- Start the Server: Open the plugin UI (through the indicator in the sidebar footer or the “Local models” button), click “Download and install server” to download the
llama-serverruntime, select a model, and then click Start. - Call the Model: Select the model provided by this plugin in the Harness configuration. The system automatically connects to
${runtimeUrl}/v1/chat/completionsto send streaming requests. - API Endpoints: The plugin provides the following HTTP endpoints:
GET /api/local-llm/models- Lists downloaded modelsPOST /api/local-llm/download- Downloads a modelGET /api/local-llm/progress- Gets download progress (SSE)POST /api/local-llm/delete- Deletes a modelGET /api/local-llm/server/status- Gets server statusPOST /api/local-llm/server/install- Installs the serverPOST /api/local-llm/server/start- Starts the serverPOST /api/local-llm/server/stop- Stops the server
Notes¶
- Environment Dependencies: Node.js 18+ is required, and Harness must include the
llmandwebServerservices. - Runtime Behavior: The
llama-serverruntime is downloaded by the plugin UI and is started only after the user clicks Start. Requests will return a clear error before the server is ready. - Directory and Port: The default server directory is
./llama-server, and the default port is8080. - Localization: Supports English, Russian, Chinese, French, Spanish, Italian, Polish, German, Hindi, and Japanese.
- Ecosystem Note: This plugin is a community plugin, and the catalog has no official affiliation with DeepSeek / High-Flyer.
Conclusion¶
dsh-local-llm provides a complete closed loop from model download to starting the inference service. For developers who need to run Llama.cpp family models locally and integrate them into the Harness ecosystem, this is a direct solution.