When developing applications based on DeepSeek Harness (DSH), upstream API rate limiting is a common pain point. Sending requests directly often leads to 429 errors or lost requests. dsh-rate-limiter is a proactive rate limiting plugin maintained by Xidong-AI, designed to control request frequency at the source.

Core Purpose

This is a preventive rate limiting plugin. It is mounted in the agent/request pipeline, executing per-provider token bucket limiting before requests are sent. When the rate limit is reached, it chooses to delay and queue requests rather than reject them or raise errors, thereby avoiding upstream 429 errors and ensuring requests are not lost.

Core Capabilities

  1. Per-provider token bucket rate limiting: Enforced before requests are sent; this is preventive rate limiting rather than remediation after failure.
  2. Queueing beyond limit: When tokens are insufficient, requests enter a delayed queue and wait for tokens to be released. This avoids directly rejecting requests and avoids producing 429 responses.
  3. Zero-intrusion design: Providers without configuration are not subject to any limits, and requests pass through directly. enabled: false can completely disable the plugin.
  4. Abort signal support: During queue waiting, if the user interrupts the operation, the wait stops immediately.
  5. No third-party dependencies: Uses a hand-written reservation-based token bucket algorithm to ensure concurrency safety, without relying on other rate limiting libraries.
  6. Coexists with retry plugin: The mount point differs from dsh-llm-retry; the two do not interfere with each other, and work better together (rate limiting first, retry second).

Installation and Configuration

Use npm to install the plugin:

dsh plugin --profile web add @xidong-ai/dsh-rate-limiter

After installation, define the rate limiting policies for each provider by configuring cordis.patch.yml. Configuration items are only supported under the providers key:

- id: rate-limiter
  config:
    enabled: true
    providers:
      nvidia:
        rate: 0.5        # 令牌填充速率(tokens/second),即长期平均 QPS
        burst: 1         # 令牌桶容量,允许的突发请求数

Parameter Description

  • rate: The number of tokens filled per second, determining the long-term average request rate.
  • burst: The capacity of the token bucket, determining how many requests can be sent consecutively before tokens are exhausted.
  • Note: Providers not listed in the configuration are not affected by rate limiting at all, and requests pass through directly.

How It Works

The plugin is mounted in the agent/request pipeline. The execution flow is as follows:

  1. Execute await next() first to retrieve the request configuration (including target provider information).
  2. Check the token bucket status based on the provider configuration.
  3. If tokens are sufficient, return the configuration immediately and send the request.
  4. If tokens are insufficient, add the request to the delayed queue.
  5. If the user interrupts the operation while waiting, the queue wait is immediately aborted by AbortSignal.

Relationship with dsh-llm-retry

The two operate at different stages of the request lifecycle and do not conflict with each other:

Plugin Intervention Time Behavior
dsh-rate-limiter Before request is sent Queued delay when limit exceeded, preventing 429 errors
dsh-llm-retry After request failure Exponential backoff retry as a fallback mechanism

Uninstallation

dsh plugin --profile web remove @xidong-ai/dsh-rate-limiter

Summary

dsh-rate-limiter solves the problem of lost requests caused by upstream API rate limiting through fine-grained token bucket control before requests are sent. Its zero-intrusion design and natural compatibility with retry plugins make it a practical tool for DSH developers handling high concurrency or restricted APIs.