Introduction

In the DeepSeek Harness (DSH) ecosystem, the official @deepseek-ai/dsh-llm-retry usually covers most request-error retry logic. However, in real deployments, relay providers sometimes mislabel transient rate-limit errors as QUOTA (insufficient balance), or a model request ends normally but returns no visible text, causing the official policy to abandon retries. As a bounded fallback recovery plugin, dsh-auto-retry fills this gap. It does not replace the official policy; instead, it takes over after the official provider policy stops retrying.

What This Is

This is a bounded fallback recovery plugin for DSH. It aims to address three specific issues:
1. Relay providers label recoverable transient 429 errors as QUOTA.
2. A model request “ends normally” but has no visible output.
3. Output is truncated by max-tokens.

The plugin intervenes when there is clear evidence, using error-signature recognition and automatic continuation mechanics, and keeps retry counts, total wait time, and auto-continuation counts all within configured budgets.

Core Features

  • Additional retry scope: After the official policy gives up, performs additional retries for RATE_LIMIT, HTTP 429, EMPTY_RESPONSE, and STREAM_CLOSED.
  • False positive detection: Retries only when relay provider error details contain indicators such as rate_limit_exceeded, rate limit, or allocated quota exceeded, avoiding pointless retries for actual balance exhaustion.
  • Automatic continuation: Handles output interruption caused by empty responses and max-tokens truncation, with at most 2 automatic continuations per turn.
  • Budget control: Limits retry counts and total wait time within configured budgets to prevent infinite retries.
  • Event reuse: Reuses DSH’s durable events (llm/retry, llm/retry-started) so session replay and existing retry-state UIs can correctly interpret subsequent actions.

Install and Enable

Install from GitHub into a Web profile:

dsh plugin --profile web add "github:Windsland52/dsh-auto-retry#main"

The installer will automatically add dependencies containing dsh.bundle.patch to the profile bundle. After installation, restart DSH.

Typical Usage

The configuration file is located in the profile’s cordis.patch.yml. Here is a basic configuration example:

- id: dsh-auto-retry
  name: '@dsh-external/dsh-auto-retry'
  config:
    # 空数组表示所有 provider;也可以只写 [scnet]
    providers: []
    # 额外重试的错误码
    retryCodes:
      - RATE_LIMIT
      - EMPTY_RESPONSE
      - STREAM_CLOSED
    # 启用预算控制
    retryQuota: true
    maxRetries: 8
    maxElapsedMs: 900000
    # 初始延迟与抖动配置
    initialDelayMs: 1500
    maxDelayMs: 60000
    jitterRatio: 0.2
    # 自动续行限制
    maxAutoContinuations: 2
    continueOnMaxTokens: true

Applicable Scenarios and Notes

  • Compatibility: Verified for DSH 0.1.0-rc.7 Web profile.
  • Behavior boundaries: The plugin does not switch models or providers, and does not modify request content. Internal automatic continuation uses user messages sourced from plugin and does not pretend to be human input.
  • Plugin exclusivity: It is not recommended to use it together with dsh-llm-bounded-retry, because both involve a global retry budget and may conflict.

Conclusion

dsh-auto-retry is positioned as a host-side, provider-aware fallback recovery tool. It preserves the official retry precedence and specifically fills the three gaps of false positive transient rate limiting from relays, empty visible output, and output truncation, providing developers with a controllable retry safety net. See its GitHub repository for more details.