Preface¶
The DSH Agent loop can issue multiple consecutive requests within tens of seconds. This can easily exceed a provider’s TPM (tokens per minute) or RPM (requests per minute) limits, causing conversations to be interrupted by 429 errors. The @zhourenke/dsh-agent-rate-limit plugin inserts a layer of adaptive latency before the LLM streaming pipeline. It uses a sliding window algorithm to manage the request queue, ordering requests based on remaining capacity before triggering the provider’s rate limit.
Plugin Overview¶
- Name:
@zhourenke/dsh-agent-rate-limit - Maintainer: zhourenke
- License: MIT
Core capabilities:
- Reads API usage data blocks, including cache hits, for accurate billing accounting.
- Delays rather than rejects requests, avoiding conversation interruptions caused by 429 errors.
- Default values align with common quotas, with no modifications required to DSH source code.
- Supports hot reloading of configuration.
Installation and Enablement¶
Install via the command line:
dsh plugin --profile web add "github:zhourenke/dsh-agent-rate-limit"
After installation, DSH must be restarted for the plugin to take effect (newly added bundles are loaded only during process startup). To confirm successful plugin loading: type /agent-rate-limit in the chat input box. If Status: loaded is displayed, the plugin is working correctly.
Configuration¶
It works with the default configuration, but parameters can be adjusted by editing the configuration file:
~/.dsh/profiles/web/cordis.patch.yml
Key configuration options include:
tpmLimit: Tokens per minute limit, default 1200000.rpmLimit: Requests per minute limit, default 15000.safetyFactor: Safety factor, default 0.8. The effective limit istpmLimit * safetyFactor.verbose: Whether to output verbose logs.countFailedAttempts: Whether failed attempts count toward the window, default true.
After editing and saving the configuration, no restart is required; hot reloading applies the changes immediately (only installing or uninstalling the plugin itself requires a restart).
How It Works and Limitations¶
The plugin controls request frequency using a sliding window algorithm. If there is sufficient capacity within the window, requests are allowed directly; when approaching the limit, it waits for older records to slide out of the window.
Notes:
- Window state is kept within the DSH process and is not shared across multiple instances.
- Configuration is global for a provider.
- Failed attempts count toward the window by default.
- It does not guarantee avoiding rate-limit violations; it only lowers the probability.
Conclusion¶
This plugin provides DSH with a basic protection mechanism by queuing requests before triggering provider rate limits, reducing the risk of conversation interruptions. For more details, see the GitHub repository.