Introduction

Running local large models usually involves tedious Python environment management, virtual environment setup, and process maintenance. For users of DeepSeek Harness (DSH), integrating a local inference service requires additional custom development. The dsh-mlx-local plugin aims to solve this pain point. It manages the Python environment, starts and stops mlx_lm.server, and seamlessly connects the local service to DSH’s existing “custom provider” system.

What this is

dsh-mlx-local is a DSH plugin designed for Apple Silicon Macs. Maintained by JshGao, its core responsibility is to provide infrastructure for local model inference. It does not involve complex business logic; instead, it focuses on environment setup, service startup and shutdown, and protocol integration with the DSH client, enabling users to access local models through DSH without incurring fixed request costs.

Core features

  • Settings page management: Provides an “MLX Models” section in DSH’s settings page, supporting starting, stopping, and switching local models.
  • Automatic environment configuration: Automatically detects Python 3.9–3.13, creates a dedicated virtual environment (venv), and installs mlx-lm.
  • Model directory management: Supports adding, removing, pre-downloading, and listing local models.
  • Service stability: Includes abnormal exit detection and cleanup of residual processes, and automatically closes the local service when DSH exits.
  • Qwen3 optimization: Automatically configures thinking strength for local Qwen3 models. The main interface model selector can switch between Off / High.
  • Clean integration: Does not register additional providers, continuing to use DSH’s “custom provider” integration; does not register system prompt sections, and does not register any tools. The plugin only runs as infrastructure.

Installation and enabling

Use the official Release package for installation:

dsh plugin --profile web add https://github.com/JshGao/dsh-mlx-local/releases/download/v0.4.1/dsh-mlx-local-0.4.1.tgz

After installation, restart DSH to load the plugin.

Typical usage

  1. Start the service: After restarting DSH, open Settings → MLX Models. Load the model directory on the settings page or use a built-in HF model, select a model, and click Start.
  2. Configure Provider: In Settings → Models → Add provider:
    • Select custom provider.
    • The route name can be arbitrary, for example local.
    • Select the openai-completions protocol.
    • Enter http://127.0.0.1:8080/v1 as the baseURL.
    • Enter the model ID displayed in the MLX plugin model directory.
    • Enter any value for the API Key (the local service does not validate it).
  3. Use a session: Create a new session and select that provider from the main interface. If the model is Qwen3, the thinking strength will display as Off / High.

Applicable scenarios and notes

  • Hardware and system: Supports only Apple Silicon (M-series) Macs, requiring macOS 13 or later.
  • DSH version: Requires DSH version 0.1.5-rc.1 or later, and the dsh command must be available.
  • Python version: Requires 3.9–3.13; 3.10–3.12 is recommended.
  • Network and storage: First use of Hugging Face models requires internet access to download weights; each 4-bit model occupies about 2–6 GB of disk space.
  • Service startup: The service does not start automatically when DSH starts; users must explicitly click Start on the settings page.
  • Version updates: 0.4.1 is a breaking change. The plugin no longer registers any tools. When upgrading from an older version, you must remove the plugin before installing the new version.
  • Uninstall leftovers: On uninstall, the plugin settings and the venv and logs under ~/.dsh/mlx/ are retained and must be deleted manually.

Closing note

dsh-mlx-local encapsulates the tedious work of MLX local inference into a standard plugin, lowering the barrier for using DSH on Apple Silicon for local deployment. For more details or to view the source code, visit the plugin directory or the GitHub repository.