Introduction¶
In the development or operations workflow of DeepSeek Harness, checking GPU resource status in real time is a frequent task. The traditional approach is to manually SSH into each server and run nvidia-smi, which is inefficient and difficult to manage centrally. The dsh-remote-gpu-monitoring plugin integrates this process into the DSH Web interface, providing a sidebar panel that displays the status of all GPU servers reachable over SSH.
What is it¶
This is a remote multi-server GPU status monitoring plugin for DeepSeek Harness Web. It reads read-only nvidia-smi information from each host over SSH, uses ControlMaster connection reuse and host-side caching, and presents a unified monitoring view in the browser.
Core capabilities¶
- Unified view across all servers: Performs fixed read-only queries for each host, displaying GPU model, memory usage, utilization, temperature, and power consumption.
- Zero configuration: Automatically parses
~/.ssh/configon startup and discovers all specific Host aliases (ignoringHost *patterns). - On-demand selection: Servers unchecked in the sidebar settings disappear from the panel, and data collection for that server stops immediately.
- Security by design: Only the Host process is responsible for opening SSH connections, using
BatchMode=yesand public key authentication, without passing user input to remote endpoints and without storing any private keys. - Low-overhead design: Reuses
ControlMastersockets, supports host-side shared caching, and minimizes refresh overhead. - Agent tools: Provides two tools:
gpu_overview(view cached snapshot) andgpu_refresh(refresh immediately).
Installation and enablement¶
Use the official installation command to add the plugin to the Web Profile:
dsh plugin --profile web add github:sshhhll002/dsh-remote-gpu-monitoring
After installation, restart dsh --profile web and refresh the Web page. The sidebar will show a “Remote GPU” entry along with the corresponding monitoring panel and tools.
Usage instructions¶
- Discover hosts: On startup, the plugin reads
~/.ssh/configand lists all specific Host aliases. - Configure selection: Click the gear icon on the sidebar panel and select the servers to monitor. Unselected servers will not appear in the panel and will not be included in data collection.
- Update configuration: After modifying
~/.ssh/configto add new hosts, you must restart dsh for the new aliases to be discovered. - Use tools: In the Agent toolbar, use
gpu_overviewto view the cached snapshot (with zero SSH overhead), or usegpu_refreshto trigger immediate collection.
Notes¶
- Prerequisites: Ensure that each GPU server has
nvidia-smiinstalled and that the system SSH client supports passwordless public key authentication (that is, thessh <alias>command does not require interactive password entry). - Host detection: If a server does not have
nvidia-smiinstalled, the plugin detects it as having no GPU and hides it, but it periodically re-probes the host. - Permissions and security: The plugin only reads
~/.ssh/configand does not store or transmit any private keys. All SSH connections are initiated by the dsh Host process.
Conclusion¶
This plugin is suitable for users who need to centrally monitor multi-node GPU resources. Through SSH connection reuse and caching mechanisms, it transforms scattered nvidia-smi checks into an efficient centralized panel. For more details, refer to the plugin directory and the source code repository.