When developing with DeepSeek Harness (DSH), parallel calls from agents or background tasks can easily trigger provider 429 errors or rate limiting. dsh-llm-rate-limit is a plugin designed to intervene before requests reach the provider, providing rate limiting, concurrency control, and queue management.

Plugin Overview

This is a DeepSeek Harness plugin maintained by Asong6824. It uses a token bucket algorithm and a FIFO queue to control the frequency, concurrency, and token consumption of LLM requests, and handles 429 errors returned by the provider.

Core Features

The plugin primarily provides the following capabilities:

  • Request and concurrency control
    Provider-scoped request control. Supports token-bucket-based requests-per-minute (RPM) limits and configurable burst capacity. It also provides concurrency limits and a bounded FIFO queue, with timeout and cancellation mechanisms.
  • Token budget management
    Supports an optional per-minute token budget. The system estimates token usage and reconciles actual usage after a successful response.
  • Adaptive cooldown
    Provides an adaptive cooldown mechanism based on provider error codes, HTTP status codes, and the Retry-After response header, avoiding continuous request submissions while the provider is rate limiting.
  • Traffic management
    Supports explicit routing of auxiliary requests, preventing background traffic from blocking primary work.
  • Lifecycle and diagnostics
    Provides persistent admission wait and start events to facilitate DSH session diagnostics. The plugin cleans up its lifecycle on disposal and does not leave behind queues or active requests.
  • Retry awareness
    Each retry attempt is treated as a separate admission decision, and the plugin itself does not perform retries.

Installation and Enablement

The installation command is as follows:

dsh plugin --profile web add dsh-llm-rate-limit

After installation, it must be enabled through the configuration file. The plugin protects the deepseek-official provider by default.

Typical Usage

Override the llm-rate-limit configuration in $DSH_HOME/profiles/<profile>/cordis.patch.yml.

Configure Provider Limits

- id: llm-rate-limit
  config:
    providers:
      deepseek-official:
        requests: { perMinute: 30, burst: 1 }
        maxConcurrentRequests: 2
        queue: { maxSize: 100, maxWaitMs: 300000, auxiliary: reject }
        cooldown:
          codes: [RATE_LIMIT, SERVER]
          statuses: [429, 529]
          initialDelayMs: 500
          maxDelayMs: 60000
          maxProviderDelayMs: 3600000
          jitterRatio: 0.1

Configure Token Budget

If you need to limit token consumption, you can add a tokens configuration:

tokens:
  perMinute: 1000000
  burst: 200000
  estimatedOutputTokens: 8192
  imageTokens: 1024

Note: tokens.burst must be large enough to accommodate the estimated usage of a complete request. You can omit the tokens configuration if token limits are not needed.

Use Cases and Notes

Use cases:
Use it when parallel agents, subagents, retry mechanisms, or background requests cause HTTP 429 errors or provider rate limiting.

Notes:
* Requirements: DeepSeek Harness 0.1.0-rc.8 or later and Node.js 22.19 or later are required.
* State persistence: The plugin state is process-local and is reset after a DSH restart.
* Functional boundaries: The plugin does not provide distributed quotas, automatic retries, or provider failover.
* License: The plugin is released under the MIT License. Review the source code before use.

Summary

dsh-llm-rate-limit provides basic traffic control for DSH. Through configuration files, it enables fine-grained control over different providers, making it suitable for scenarios that require managing request frequency and concurrency within a single process.

Plugin Directory | GitHub Repository