AI Agent Hub
Back to skills
LLM Key Pool Tiered Polling Management icon

LLM Key Pool Tiered Polling Management

AI Agent Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_7c8b7398/llm-key-pool.

About this skill

Problem

When an Agent needs stable access to large-model APIs, a single provider can fail because of quota limits, 429 rate limiting, expired keys, or a single point of failure. llm_key_pool organizes API keys from multiple providers into a configurable tiered pool, allowing callers to use a unified OpenAI-compatible interface while the pool handles throttling and failover. This is especially relevant for batch inference, parallel multi-Agent workloads, and long-running automation flows.

How It Works and Limits

The skill uses llm_config.yaml to define three tiers: a primary tier for high-quota providers such as Alibaba Cloud Bailian and Zhipu AI, a daily refresh tier for daily-reset providers such as Volcengine and Google AI Studio, and a fallback tier for open-source or aggregated providers such as SiliconFlow and OpenRouter. The workflow is: select a primary-tier key for the request; on a 429 response, mark the key as cooling down and switch to the next usable key or the next tier; restore the key automatically after cooldown. --status reports key-pool usage across tiers, and --test validates the configuration.

Keep in mind that the config file contains sensitive API keys and should not be committed to version control. Cross-tier fallback adds a small amount of latency. Configure at least one usable provider per tier, and prefer platforms that expose OpenAI-compatible endpoints.

Use Cases

  • Prevent batch LLM generation from stopping when one provider returns 429 rate-limit errors.
  • Route parallel multi-Agent model calls through tiered API keys that switch to available quota.
  • Keep long automation flows running when the primary tier exhausts its quota during execution.
  • Validate multi-provider key configuration and inspect pool status before deployment.

Best For

  • Agent infrastructure engineers who need one interface to multiple model providers and throttling handling.
  • Backend engineers building batch inference scripts who want 429 cooldown and automatic key switching.
  • Platform engineers setting up automation pipelines who need primary, daily, and fallback quota tiers.
  • LLM gateway maintainers who need to inspect key-pool status and validate provider configuration.