When developing with DeepSeek Harness (DSH), parallel calls from agents or background tasks can easily trigger provider 429 errors or rate limiting. dsh-llm-rate-limit is a plugin designed to intervene before requests reach the provider, providing rate limiting, concurrency control, and queue management.
Plugin Overview¶
This is a DeepSeek Harness plugin maintained by Asong6824. It uses a token bucket algorithm and a FIFO queue to control the frequency, concurrency, and token consumption of LLM requests, and handles 429 errors returned by the provider.
Core Features¶
The plugin primarily provides the following capabilities:
- Request and concurrency control
Provider-scoped request control. Supports token-bucket-based requests-per-minute (RPM) limits and configurable burst capacity. It also provides concurrency limits and a bounded FIFO queue, with timeout and cancellation mechanisms. - Token budget management
Supports an optional per-minute token budget. The system estimates token usage and reconciles actual usage after a successful response. - Adaptive cooldown
Provides an adaptive cooldown mechanism based on provider error codes, HTTP status codes, and theRetry-Afterresponse header, avoiding continuous request submissions while the provider is rate limiting. - Traffic management
Supports explicit routing of auxiliary requests, preventing background traffic from blocking primary work. - Lifecycle and diagnostics
Provides persistent admission wait and start events to facilitate DSH session diagnostics. The plugin cleans up its lifecycle on disposal and does not leave behind queues or active requests. - Retry awareness
Each retry attempt is treated as a separate admission decision, and the plugin itself does not perform retries.
Installation and Enablement¶
The installation command is as follows:
dsh plugin --profile web add dsh-llm-rate-limit
After installation, it must be enabled through the configuration file. The plugin protects the deepseek-official provider by default.
Typical Usage¶
Override the llm-rate-limit configuration in $DSH_HOME/profiles/<profile>/cordis.patch.yml.
Configure Provider Limits¶
- id: llm-rate-limit
config:
providers:
deepseek-official:
requests: { perMinute: 30, burst: 1 }
maxConcurrentRequests: 2
queue: { maxSize: 100, maxWaitMs: 300000, auxiliary: reject }
cooldown:
codes: [RATE_LIMIT, SERVER]
statuses: [429, 529]
initialDelayMs: 500
maxDelayMs: 60000
maxProviderDelayMs: 3600000
jitterRatio: 0.1
Configure Token Budget¶
If you need to limit token consumption, you can add a tokens configuration:
tokens:
perMinute: 1000000
burst: 200000
estimatedOutputTokens: 8192
imageTokens: 1024
Note: tokens.burst must be large enough to accommodate the estimated usage of a complete request. You can omit the tokens configuration if token limits are not needed.
Use Cases and Notes¶
Use cases:
Use it when parallel agents, subagents, retry mechanisms, or background requests cause HTTP 429 errors or provider rate limiting.
Notes:
* Requirements: DeepSeek Harness 0.1.0-rc.8 or later and Node.js 22.19 or later are required.
* State persistence: The plugin state is process-local and is reset after a DSH restart.
* Functional boundaries: The plugin does not provide distributed quotas, automatic retries, or provider failover.
* License: The plugin is released under the MIT License. Review the source code before use.
Summary¶
dsh-llm-rate-limit provides basic traffic control for DSH. Through configuration files, it enables fine-grained control over different providers, making it suitable for scenarios that require managing request frequency and concurrency within a single process.