When deploying local large model inference in a WSL2 environment, GPU visibility and VRAM usage are often central to troubleshooting. When Ollama, llama-server, or vLLM are running at the same time, the risk of VRAM conflicts (OOM) is high. The dsh-wsl-gpu plugin for DeepSeek Harness (DSH) is designed to address GPU status monitoring and port conflict detection in this scenario.

Plugin Positioning

dsh-wsl-gpu is a DeepSeek Harness plugin that belongs to the dsh-wsl-kit toolkit. By parsing nvidia-smi output, it provides detailed GPU status, Blackwell architecture hints, and inference port occupancy, helping developers quickly identify resource bottlenecks in local inference environments.

Core Features

The plugin provides the following capabilities:

  1. GPU status parsing: Extracts GPU name, driver version, VRAM used and total, GPU utilization, and compute capability.
  2. New architecture hints: Provides specific hints for Blackwell / RTX 50 series (sm_120) and CUDA 12.8+/13.x environments.
  3. Port scanning: Automatically scans common inference service ports (11434, 1234, 8000, 8080) to determine port occupancy.
  4. Troubleshooting guidance: Provides pointers to host_reach and docker_doctor focus=vllm to assist further investigation.

Installation and Enablement

The plugin can be installed using either of the following methods.

Method 1: Install via Kit (Recommended)

Use the dsh-wsl-kit installation script, specifying KIT_SET=llm to include this plugin:

curl -fsSL https://raw.githubusercontent.com/173787247/dsh-wsl-kit/master/install.sh | KIT_SET=llm bash

Method 2: Install via DSH Plugin Manager

Directly add it using the DSH plugin command:

dsh plugin --profile web add github:173787247/dsh-wsl-gpu

Typical Usage

It is recommended to run gpu_doctor in the following scenarios:

  • After driver updates: Verify the visibility of Windows NVIDIA drivers in WSL.
  • Before build failures: When encountering a CUDA build failure, quickly check GPU status.
  • Before loading models: Before loading large GGUF models, check VRAM pressure and port occupancy to avoid out-of-memory issues.

Notes and Compatibility

  • Toolkit membership: This plugin is part of dsh-wsl-kit. It is recommended to use the KIT_SET=daily, llm, github, or full installation sets (see the kit README for details).
  • DSH version requirement: DeepSeek Harness (dsh) version >= 0.1.2 is required.
  • Cloud Flash configuration: The plugin itself does not configure the Cloud Flash model. To use Cloud Flash (V4.1 Flash), manually set the model ID to deepseek-flash in ~/.dsh/settings.yaml or llm-deepseek.
  • Agent Teams: The Agent Teams feature is an experimental upstream feature; this plugin does not depend on it.
  • Version information: The current plugin version is 0.2.2 (full version; also included in llm). The latest verified version in the kit is 0.2.0-rc.2.

This tool focuses on GPU status diagnostics in WSL environments and is suitable for developers who need fine-grained management of local inference resources. For more details, please visit the GitHub repository.