When deploying local large model inference in a WSL2 environment, GPU visibility and VRAM usage are often central to troubleshooting. When Ollama, llama-server, or vLLM are running at the same time, the risk of VRAM conflicts (OOM) is high. The dsh-wsl-gpu plugin for DeepSeek Harness (DSH) is designed to address GPU status monitoring and port conflict detection in this scenario.
Plugin Positioning¶
dsh-wsl-gpu is a DeepSeek Harness plugin that belongs to the dsh-wsl-kit toolkit. By parsing nvidia-smi output, it provides detailed GPU status, Blackwell architecture hints, and inference port occupancy, helping developers quickly identify resource bottlenecks in local inference environments.
Core Features¶
The plugin provides the following capabilities:
- GPU status parsing: Extracts GPU name, driver version, VRAM used and total, GPU utilization, and compute capability.
- New architecture hints: Provides specific hints for Blackwell / RTX 50 series (
sm_120) and CUDA 12.8+/13.x environments. - Port scanning: Automatically scans common inference service ports (
11434,1234,8000,8080) to determine port occupancy. - Troubleshooting guidance: Provides pointers to
host_reachanddocker_doctor focus=vllmto assist further investigation.
Installation and Enablement¶
The plugin can be installed using either of the following methods.
Method 1: Install via Kit (Recommended)
Use the dsh-wsl-kit installation script, specifying KIT_SET=llm to include this plugin:
curl -fsSL https://raw.githubusercontent.com/173787247/dsh-wsl-kit/master/install.sh | KIT_SET=llm bash
Method 2: Install via DSH Plugin Manager
Directly add it using the DSH plugin command:
dsh plugin --profile web add github:173787247/dsh-wsl-gpu
Typical Usage¶
It is recommended to run gpu_doctor in the following scenarios:
- After driver updates: Verify the visibility of Windows NVIDIA drivers in WSL.
- Before build failures: When encountering a CUDA build failure, quickly check GPU status.
- Before loading models: Before loading large GGUF models, check VRAM pressure and port occupancy to avoid out-of-memory issues.
Notes and Compatibility¶
- Toolkit membership: This plugin is part of
dsh-wsl-kit. It is recommended to use theKIT_SET=daily,llm,github, orfullinstallation sets (see the kit README for details). - DSH version requirement: DeepSeek Harness (
dsh) version>= 0.1.2is required. - Cloud Flash configuration: The plugin itself does not configure the Cloud Flash model. To use Cloud Flash (V4.1 Flash), manually set the model ID to
deepseek-flashin~/.dsh/settings.yamlorllm-deepseek. - Agent Teams: The Agent Teams feature is an experimental upstream feature; this plugin does not depend on it.
- Version information: The current plugin version is 0.2.2 (full version; also included in llm). The latest verified version in the kit is 0.2.0-rc.2.
This tool focuses on GPU status diagnostics in WSL environments and is suitable for developers who need fine-grained management of local inference resources. For more details, please visit the GitHub repository.