Introduction¶
When using shared LLM gateways (such as campus network proxies, self-hosted gateways, or provider pools), the gateway often returns a 429 status code along with an explicit retry_after_seconds hint in the response body, for example: “All upstream providers are cooling down. Please retry after 28 seconds.”
The built-in llm-retry mechanism in DeepSeek Harness supports failure.providerRetryAfterMs, but this field is usually populated by adapters parsing HTTP headers. When the hint exists only in the response body text (as with the llm-pi-ai adapter, which does not parse headers), llm-retry falls back to a local backoff strategy (500ms → 1s, 2 attempts total), causing the request to fail within the cooldown window, even though it would succeed after waiting 28 seconds. This plugin aims to fill that gap.
Plugin Overview¶
dsh-cooldown-retry is a DeepSeek Harness plugin that provides a patient automatic retry mechanism. It reads the retry delay from upstream 429 capacity cooldown messages instead of giving up after two quick attempts.
- Owner: dboycht
- Category: Workflow
- License: MIT
Core Features¶
- Patient retries: Automatically retries DeepSeek Harness instead of giving up after two quick attempts.
- Reads upstream delays: Reads retry delays (such as
retry_after_seconds) from upstream 429 capacity cooldown messages. - Absorbs mislabeled 400s: Absorbs 400 errors mislabeled as “audio modality not supported” when the gateway is saturated.
- Budget management: Manages retry budgets per turn and per provider.
- Cancellation-aware waiting: Supports ending the wait immediately when the turn is cancelled.
Installation and Activation¶
We recommend adding the plugin using the official installation command.
dsh plugin --profile web add github:dboycht/dsh-cooldown-retry
Prerequisite: The installation command requires git to be available in the system PATH (for the github: specifier).
After installation, reload the configuration for the change to take effect; no process restart is required.
Verify that the installation succeeded:
dsh --profile web --dump-config | grep -A2 'cooldown-retry'
Note: The plugin entry must appear after the built-in llm-retry entry. agent/request-error is a waterfall pipeline; later-registered listeners act as inner listeners, so this plugin can intercept failures first.
Configuration and Usage¶
The default configuration is:
* maxRetries: 5
* minDelayMs: 1000
* maxDelayMs: 300000
* acrossSteps: true
* mislabeledClientError: true
You can override the defaults with custom configuration. Because non-insert patches replace the entire configuration line, all keys that should be preserved must be restated when making changes.
- id: cooldown-retry
config:
maxRetries: 10
maxDelayMs: 600000
Key configuration options:
* maxRetries: The number of patient retries per turn/per provider before the failure is passed downstream.
* minDelayMs / maxDelayMs: The window bounds for each wait (1s - 300s).
* acrossSteps: Defaults to true, meaning a single retry budget is shared across all steps in a turn; set to false to restore independent budgets per step.
* mislabeledClientError: Defaults to true. If the same provider has recently experienced capacity cooldown and the gateway returns a 400 error claiming the modality is “not supported”, the plugin treats it as a circuit replay and waits instead of immediately raising an error.
* mislabeledMaxRetries: A separate retry budget used specifically for “mislabeled 400” handling, independent of maxRetries.
* mislabeledEvidenceMs: The validity window in which a capacity cooldown counts as “evidence”.
Limitations:
* maxRetries and mislabeledMaxRetries are capped at 100.
* Non-insert patches replace the entire configuration line, so all keys that should be retained must be restated.
Observability¶
All waits and abandoned retries are logged through ctx.logger with the label cooldown-retry. This helps trace specific retry behavior and delay durations in the logs.
Conclusion¶
By parsing cooldown hints in the upstream response body, this plugin addresses a common 429 handling issue in shared gateway environments that is ignored by built-in retry policies, and it provides fault tolerance for gateway errors that are mislabeled. For DSH workflows that rely on unstable gateways or require high availability, this is a practical enhancement.