Foreword

In the agent development workflow of DeepSeek Harness (DSH), unpredictable test results are a common pain point. Agents need to distinguish between “code logic errors” and “instability in the tests themselves (Flaky Tests).” dsh-flakefinder is a workflow plugin that repeatedly runs tests to identify unstable test cases, provides historical records and a quarantine list, and helps Agents track test stability.

Plugin Overview

Name: dsh-flakefinder

Positioning: DeepSeek Harness test stability plugin

Maintainer: STARDUSTLC666

License: MIT

Core Value: zero runtime dependencies, support for multiple testing frameworks, and management of flaky test cases through a quarantine list.

Core Features

The plugin provides five core tools, covering the complete workflow from detection to quarantine:

  • flaky_detect: runs a test N times repeatedly and determines the result as stable-pass / stable-fail / flaky.
  • flaky_history: queries detection history and filters by target.
  • flaky_report: aggregates the quarantine list and historical data to output a stability report.
  • flaky_quarantine: writes flaky test cases into the quarantine list (write operation).
  • flaky_clear: removes recovered test cases from the quarantine list (write operation).

Installation and Activation

Use the official installation command to enable the plugin:

dsh plugin --profile web add dsh-flakefinder

Typical Usage

Agents can handle unstable tests through interactive instructions:

User: src/checkout.test.ts has been failing recently. Help me determine the cause.
Agent:
  flaky_detect(target="src/checkout.test.ts", runs=5)
  → Determination: flaky; 3/5 passed, failures are concentrated in the useFakeTimers test case
  → flaky_quarantine(tests=["src/checkout.test.ts > 定时器恢复"], reason="定时器竞态")

Configuration

The plugin supports behavior customization through configuration options. The main options are as follows:

  • defaultRuns: the default number of repeated runs for flaky_detect, default 5.
  • maxRuns: the maximum number of repeated runs allowed per invocation, default 20.
  • timeoutMs: the timeout for a single test run, default 120000.
  • graceMs: the grace period after timeout, default 10000.
  • writeApproval: whether writes to the quarantine list go through an approval gate, default true.
  • dataDir: the directory for history storage, default DSH_HOME/.dsh-flakefinder.
  • quarantineFile: the path to the quarantine list, default <cwd>/.flakefinder.json.
  • pythonPath: the Python interpreter used by pytest, default DSH_FLAKEFINDER_PYTHON or python/python3.

Technical Details and Notes

  • Zero runtime dependencies: The plugin itself does not depend on external runtime libraries.
  • Quarantine list mechanism: The list only records entries and does not modify test source code. Agents should consult flaky_report before performing quarantine.
  • Process service: Test processes use the official DSH subprocess service, are invoked with an argv array, and do not rely on a shell.
  • Framework support: Supports vitest, jest, pytest, and node:test, with automatic detection or explicit framework specification.
  • Compatibility requirement: Requires Node.js 22.19 or later (22.x or 24.x).
  • Uninstall note: The Web service must be restarted after uninstalling for the change to take effect.

Summary

dsh-flakefinder provides DSH environments with a standard toolchain for identifying and managing unstable tests. Use flaky_detect to locate issues, flaky_quarantine to isolate them, and flaky_history to track history, which can effectively improve the efficiency of assessing test environment stability.

GitHub Repository