Introduction

In the development or operations workflow of DeepSeek Harness, checking GPU resource status in real time is a frequent task. The traditional approach is to manually SSH into each server and run nvidia-smi, which is inefficient and difficult to manage centrally. The dsh-remote-gpu-monitoring plugin integrates this process into the DSH Web interface, providing a sidebar panel that displays the status of all GPU servers reachable over SSH.

What is it

This is a remote multi-server GPU status monitoring plugin for DeepSeek Harness Web. It reads read-only nvidia-smi information from each host over SSH, uses ControlMaster connection reuse and host-side caching, and presents a unified monitoring view in the browser.

Core capabilities

  • Unified view across all servers: Performs fixed read-only queries for each host, displaying GPU model, memory usage, utilization, temperature, and power consumption.
  • Zero configuration: Automatically parses ~/.ssh/config on startup and discovers all specific Host aliases (ignoring Host * patterns).
  • On-demand selection: Servers unchecked in the sidebar settings disappear from the panel, and data collection for that server stops immediately.
  • Security by design: Only the Host process is responsible for opening SSH connections, using BatchMode=yes and public key authentication, without passing user input to remote endpoints and without storing any private keys.
  • Low-overhead design: Reuses ControlMaster sockets, supports host-side shared caching, and minimizes refresh overhead.
  • Agent tools: Provides two tools: gpu_overview (view cached snapshot) and gpu_refresh (refresh immediately).

Installation and enablement

Use the official installation command to add the plugin to the Web Profile:

dsh plugin --profile web add github:sshhhll002/dsh-remote-gpu-monitoring

After installation, restart dsh --profile web and refresh the Web page. The sidebar will show a “Remote GPU” entry along with the corresponding monitoring panel and tools.

Usage instructions

  1. Discover hosts: On startup, the plugin reads ~/.ssh/config and lists all specific Host aliases.
  2. Configure selection: Click the gear icon on the sidebar panel and select the servers to monitor. Unselected servers will not appear in the panel and will not be included in data collection.
  3. Update configuration: After modifying ~/.ssh/config to add new hosts, you must restart dsh for the new aliases to be discovered.
  4. Use tools: In the Agent toolbar, use gpu_overview to view the cached snapshot (with zero SSH overhead), or use gpu_refresh to trigger immediate collection.

Notes

  • Prerequisites: Ensure that each GPU server has nvidia-smi installed and that the system SSH client supports passwordless public key authentication (that is, the ssh <alias> command does not require interactive password entry).
  • Host detection: If a server does not have nvidia-smi installed, the plugin detects it as having no GPU and hides it, but it periodically re-probes the host.
  • Permissions and security: The plugin only reads ~/.ssh/config and does not store or transmit any private keys. All SSH connections are initiated by the dsh Host process.

Conclusion

This plugin is suitable for users who need to centrally monitor multi-node GPU resources. Through SSH connection reuse and caching mechanisms, it transforms scattered nvidia-smi checks into an efficient centralized panel. For more details, refer to the plugin directory and the source code repository.