Preface

When developing model inference with DeepSeek Harness, encountering rate limit (429) or quota exhaustion errors returned by model providers is common. Manually switching or retrying is not only inefficient but also prone to interrupting the user experience. The dsh-llm-failover plugin aims to solve this problem by automatically detecting errors and switching to the next provider, ensuring the continuity of inference services.

What It Is

dsh-llm-failover is a DeepSeek Harness (DSH) plugin maintained by HB00. It implements failover for model providers. When the current provider fails due to rate limiting or quota issues, the plugin automatically switches to the next provider based on the configuration until it finds an available service or reaches the permanent fallback provider.

Core Features

  • Automatic Failover: Automatically switches providers when a RATE_LIMIT or QUOTA error is detected.
  • Per-Provider Model Configuration: After switching, the new provider can use a different model.
  • Cooldown Mechanism: Providers are skipped during cooldown to avoid frequent switching until the cooldown expires.
  • Permanent Fallback: The last provider in the configured list acts as the permanent fallback and is never switched away from.
  • UI Notification: Upon each switch, a disappearing notification bar appears above the input field in the interface.
  • Multiple Configuration Methods: Supports configuration via a settings card or a YAML file.

Installation

Add the plugin using the official installation command:

dsh plugin --profile web add dsh-llm-failover

Configuration and Usage

The plugin configuration file is located at ~/.dsh/settings.yaml. There are two ways to configure it: fill it in through the “Plugin Configuration” card in the settings interface, or edit the YAML file directly.

YAML Configuration Example

llm-failover:
  enabled: true
  providers:
    - provider: huoshan
      model: deepseek-v4-flash
    - provider: huoshan2
      model: deepseek-v4-flash
    - provider: deepseek-official
      model: deepseek-v4-flash
  fallbackAfterRetries: 2
  cooldownMs: 60000

Configuration Item Description:
* enabled: Master switch that controls whether the plugin is active.
* providers: List of providers, attempted in order. The provider at the end of the list is the permanent fallback.
* fallbackAfterRetries: The number of consecutive failures after which failover and cooldown are triggered.
* cooldownMs: Cooldown duration for a provider (in milliseconds).

How It Works

The plugin implements its logic by hooking into two official Agent hooks in DSH:
1. agent/request: On request, selects the first non-cooldown provider from the configured list.
2. agent/request-error: On error, determines whether to trigger failover based on fallbackAfterRetries. If triggered, it puts the current provider into a cooldown state and returns a retry instruction. The provider at the end of the list is never switched away from.

Notes

  • Dependencies: The plugin depends on zod, @deepseek-ai/cordis, @deepseek-ai/dsh-settings, and @deepseek-ai/schemastery.
  • Permissions and Source Code: The plugin runs with the permissions of the current DSH process. It is recommended to review the source code and license (MIT) before installation.
  • Directory Affiliation: This plugin is listed in the community directory and has no official relationship with DeepSeek or High-Flyer.

Summary

dsh-llm-failover provides a simple and efficient solution that enables automatic fault tolerance across multiple model providers through configuration, making it suitable for DeepSeek Harness agent development scenarios requiring high availability.