Introduction

DeepSeek Harness (DSH) adopts an “everything is a plugin” ecosystem philosophy. When selecting and deploying LLM models, key metrics to consider include memory capacity, precision configuration, interconnect bandwidth, time to first token (TTFT), throughput, and power consumption. The dsh-model-deploy plugin aims to address pre-deployment analysis by producing an auditable estimation report based on a given model, hardware, and parameters.

Plugin Overview

dsh-model-deploy is an LLM model selection and deployment analysis tool maintained by lhwwxy under the MIT open source license. It covers combination analysis of 38 models and 20 GPU/NPU devices, and can output deployment feasibility, resource breakdowns, performance metrics, and power consumption estimates.

Installation and Activation

Installing this plugin requires specifying the --profile parameter.

dsh plugin --profile web add github:lhwwxy/dsh-model-deploy

After installation, the tool schema does not take effect immediately. You need to restart the session or reload the configuration file.

Core Features

This plugin provides the following analytical capabilities:

  • Deployment checking and remediation: Checks whether memory, precision, and interconnect requirements are met, and provides remediation suggestions.
  • Resource breakdown: Outputs memory breakdowns for single-card and cluster configurations (weights, KV cache, activations, runtime).
  • Latency analysis: Analyzes time to first token (TTFT) and overall latency under idle and fully loaded conditions.
  • Throughput analysis: Computes tok/s and req/s per request and for the system, as well as prefill throughput.
  • Power consumption estimation: Estimates idle, typical, and peak power consumption, including PUE and energy per token.
  • Parallel strategy selection: Automatically calculates and provides the rationale for selecting a TP×PP×DP parallel strategy.
  • Coverage: Supports analysis of 38 models (such as Qwen3, Llama, DeepSeek-V3, GLM, etc.) and 20 hardware devices (NVIDIA H100/H200, Huawei Ascend 910B/910C/950PR, etc.).

Typical Usage

The tool is divided into two parts: model_deploy_catalog and model_deploy_analyze.

  1. Query resources: Use model_deploy_catalog to query supported model IDs and GPU/NPU IDs.
  2. Run analysis: Invoke the model_deploy_analyze tool with parameters to perform the analysis.

Sample invocation is as follows:

{
  "model": "deepseek-v3",
  "gpu": "ascend910b",
  "gpusPerNode": 8,
  "nodes": 1,
  "precision": "int4",
  "ctx": 131072,
  "batch": 4
}

Parameter descriptions:
* model: Catalog ID or custom model JSON string.
* gpu: Catalog ID.
* gpusPerNode: Number of GPUs per node (1–8).
* nodes: Number of nodes (1–64).
* precision: Precision, supporting fp32|bf16|fp16|fp8|int8|int4|fp4.
* ctx/batch/prompt/output: Corresponding token counts.
* intraMode/intraGBs: Intra-node interconnect mode and bandwidth.
* interMode/interGBs: Inter-node interconnect mode and bandwidth.

Notes

  • Installation requirements: The --profile parameter must be provided during installation, and the session or configuration file must be restarted/reloaded after installation for changes to take effect.
  • Estimation model: The formulas and coefficients provided by the tool are public, and conclusions can be verified. It has been validated against real deployment anchors (e.g., 70B BF16 on 8xH100 ≈ 150 tok/s).
  • NPU practical differences: The actual utilization of the software stack for domestic NPUs may differ from ideal estimates, and the tool indicates this. It is recommended to perform measurements using the target framework before formal procurement.
  • Ecosystem context: The DSH philosophy is “everything is a plugin,” and this plugin has no official affiliation with DeepSeek or High-Flyer.

Summary

dsh-model-deploy provides end-to-end deployment analysis capabilities from memory to power consumption, making it suitable for rapid validation and capacity planning before deploying models. For the complete model list and hardware specifications, refer to the GitHub repository.