The DeepSeek Harness (DSH) philosophy is “everything is a plugin.” In agent development, long-running tasks spawned by ctx.subprocess.spawn (such as background builds and dependency installation) can be hard to detect when they stall. Existing monitoring approaches either rely on complex process-table queries or risk killing tasks that are still running. dsh-stall-sentinel is a lightweight watchdog plugin designed specifically for DSH profile bundles. It takes over process spawning logic and uses probe sampling to provide stall warnings and forensic reports, without interfering with task execution.

Plugin Overview

This is a pure JavaScript Cordis plugin with zero additional runtime dependencies (using only built-in Node modules). It reworks the idea behind longtask-stall-guard, implementing it from scratch for DSH scenarios. It is intended to provide automatic detection and forensics for stalled processes, and it never performs kill, terminate, or retry operations.

Core Capabilities

  • Zero additional runtime dependencies: It depends only on node:* built-in modules and does not introduce complex external libraries.
  • Automatic monitoring and wiring: By default, after installation, it automatically intercepts ctx.subprocess.spawn. You do not need to manually call guard.watch to cover bash tools, background pwsh, and the build tasks they launch.
  • Conservative decision model: It uses dual-signal forensic aggregation based on “progress freeze” and “CPU state (idle or saturated).” A stall is only declared after two consecutive sampling windows satisfy the conditions.
  • Alert and forensics only: The only handling action is “alert + generate a forensic report.” It never performs kill, terminate, or retry operations.
  • Complete logging: It supports JSONL event logging and rotation for post-incident analysis.

Installation

The plugin is provided as a dsh.bundle.patch package. During installation, make sure it is ordered after the base package.

dsh plugin --profile <profile> add <path-or-link-to-this-dir>

Or reference the local path directly in the profile’s cordis.patch.yml:

- insert:
    - id: stall-sentinel
      name: './plugins/dsh-stall-sentinel/lib/index.js'
      config: { enabled: true }

Usage

The plugin exposes a control API through ctx.provide('stall-sentinel', api). You can manually register a target that should be monitored:

const guard = ctx.get('stall-sentinel');
const { taskId } = guard.watch({
  taskId: 'my-build',
  label: 'pnpm install',
  pid: child.pid,
  logPath: 'C:\\agent\\install.log',
  marker: 'added'
});
guard.getStatus();
guard.list();
guard.unwatch(taskId);

Important: A target must provide logPath (used for the progress signal) and pid. A target without logPath will remain in the unknown state permanently, making it impossible to distinguish between slow execution and a stall.

Applicable Scenarios and Notes

  • Applicable scenarios: Applicable to long-running tasks started through ctx.subprocess.spawn, such as background pwsh, npm install, and build scripts. These tasks are automatically monitored by default.
  • Boundary notes:
  • For pure job records in ctx.jobs (without pid and without exposed output), the plugin only tracks liveness and cannot determine whether the work is stalled.
  • At the ctx.shell level, the plugin covers the behavior through the subprocess seam and does not directly consume output.
  • The plugin can still work in restricted sandboxes (where Node pipe spawning is disabled) by using stdio: ['ignore', fdOut, fdErr].
  • Reliability contract: The plugin never kills processes; it only reads probes. Sampling failures degrade to unknown. Log persistence errors fall back to console.error.