Introduction

The core philosophy of DeepSeek Harness (DSH) is “everything is a plugin.” For developers or users who need to deploy local model inference, managing process startup, VRAM usage, and parameter switching is a common pain point. The dsh-plugin-local-model plugin directly solves these problems by independently managing local models in the settings page, enabling automation and resource release.

Plugin Positioning

This is a DeepSeek Harness plugin maintained by DZQJOKER and falls under the model inference category.

Core value:
An independent “Local Model” page in the settings automatically launches llama.cpp on the first conversation, and automatically unloads the model to release resources after 5 minutes of idle time.

Core Features

Model and Process Management

  • Automatic startup: After a user configures a local GGUF model in the settings page, the first conversation will automatically launch llama.cpp and load the model.
  • Automatic unloading: If there is no interaction for 5 consecutive minutes, the plugin will automatically unload the model and release VRAM, reducing system resource usage.

Parameter Presets and Switching

The top of the settings page includes a parameter preset area. Users can save a full set of loading/inference parameters, including the selected model and vision projection, as named entries. Multiple preset groups are supported, and they can be renamed, overwritten, or deleted. Switching can be done with a single click.

Environment and Permission Isolation

  • Offline operation: The plugin does not download models over the network.
  • File safety: It does not touch workspace files.
  • Network listening: It does not listen on any ports except the local loopback address.

Typical Usage

  1. Prepare the environment: Users need to download the required GGUF model files and the llama tools themselves, then place them in the directories specified by the plugin.
  2. Configure the model: Find the “Local Model” page in DSH settings and add the configured model.
  3. Configure parameters: In the parameter preset area at the top of the settings page, save the current parameter configuration, including the model path, inference parameters, and other settings.
  4. Start the conversation: Start the first conversation; llama.cpp will automatically start and load the model.
  5. Release resources: After the conversation ends, stop input. After 5 minutes without interaction, the model will be automatically unloaded.

Installation and Enablement

An official one-click installation command is not currently provided. It is recommended to install manually through the plugin directory or build from source. Refer to the GitHub repository for detailed installation steps.

Use Cases and Notes

  • Use cases: Scenarios requiring local large model execution, strong privacy protection with no network access, or frequent switching between different inference parameters.
  • Notes:
    • Users must download the model files and llama tools themselves and place them in the specified directories.
    • The plugin runs with the permissions of the current DSH process. Before installation, check the source code and license (MIT).

Conclusion

By automating model startup and unloading, this plugin lowers the barrier to using local model inference, while parameter presets improve configuration efficiency. For developers who pursue localized and private deployment, this is a practical tool.