Preface

The DSH Agent loop can issue multiple consecutive requests within tens of seconds. This can easily exceed a provider’s TPM (tokens per minute) or RPM (requests per minute) limits, causing conversations to be interrupted by 429 errors. The @zhourenke/dsh-agent-rate-limit plugin inserts a layer of adaptive latency before the LLM streaming pipeline. It uses a sliding window algorithm to manage the request queue, ordering requests based on remaining capacity before triggering the provider’s rate limit.

Plugin Overview

  • Name: @zhourenke/dsh-agent-rate-limit
  • Maintainer: zhourenke
  • License: MIT

Core capabilities:

  • Reads API usage data blocks, including cache hits, for accurate billing accounting.
  • Delays rather than rejects requests, avoiding conversation interruptions caused by 429 errors.
  • Default values align with common quotas, with no modifications required to DSH source code.
  • Supports hot reloading of configuration.

Installation and Enablement

Install via the command line:

dsh plugin --profile web add "github:zhourenke/dsh-agent-rate-limit"

After installation, DSH must be restarted for the plugin to take effect (newly added bundles are loaded only during process startup). To confirm successful plugin loading: type /agent-rate-limit in the chat input box. If Status: loaded is displayed, the plugin is working correctly.

Configuration

It works with the default configuration, but parameters can be adjusted by editing the configuration file:

~/.dsh/profiles/web/cordis.patch.yml

Key configuration options include:

  • tpmLimit: Tokens per minute limit, default 1200000.
  • rpmLimit: Requests per minute limit, default 15000.
  • safetyFactor: Safety factor, default 0.8. The effective limit is tpmLimit * safetyFactor.
  • verbose: Whether to output verbose logs.
  • countFailedAttempts: Whether failed attempts count toward the window, default true.

After editing and saving the configuration, no restart is required; hot reloading applies the changes immediately (only installing or uninstalling the plugin itself requires a restart).

How It Works and Limitations

The plugin controls request frequency using a sliding window algorithm. If there is sufficient capacity within the window, requests are allowed directly; when approaching the limit, it waits for older records to slide out of the window.

Notes:

  • Window state is kept within the DSH process and is not shared across multiple instances.
  • Configuration is global for a provider.
  • Failed attempts count toward the window by default.
  • It does not guarantee avoiding rate-limit violations; it only lowers the probability.

Conclusion

This plugin provides DSH with a basic protection mechanism by queuing requests before triggering provider rate limits, reducing the risk of conversation interruptions. For more details, see the GitHub repository.