LLM Key Pool Tiered Polling Management
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @user_7c8b7398/llm-key-pool.
About this skill
Problem
When an Agent needs stable access to large-model APIs, a single provider can fail because of quota limits, 429 rate limiting, expired keys, or a single point of failure. llm_key_pool organizes API keys from multiple providers into a configurable tiered pool, allowing callers to use a unified OpenAI-compatible interface while the pool handles throttling and failover. This is especially relevant for batch inference, parallel multi-Agent workloads, and long-running automation flows.
How It Works and Limits
The skill uses llm_config.yaml to define three tiers: a primary tier for high-quota providers such as Alibaba Cloud Bailian and Zhipu AI, a daily refresh tier for daily-reset providers such as Volcengine and Google AI Studio, and a fallback tier for open-source or aggregated providers such as SiliconFlow and OpenRouter. The workflow is: select a primary-tier key for the request; on a 429 response, mark the key as cooling down and switch to the next usable key or the next tier; restore the key automatically after cooldown. --status reports key-pool usage across tiers, and --test validates the configuration.
Keep in mind that the config file contains sensitive API keys and should not be committed to version control. Cross-tier fallback adds a small amount of latency. Configure at least one usable provider per tier, and prefer platforms that expose OpenAI-compatible endpoints.
Use Cases
- Prevent batch LLM generation from stopping when one provider returns 429 rate-limit errors.
- Route parallel multi-Agent model calls through tiered API keys that switch to available quota.
- Keep long automation flows running when the primary tier exhausts its quota during execution.
- Validate multi-provider key configuration and inspect pool status before deployment.
Best For
- Agent infrastructure engineers who need one interface to multiple model providers and throttling handling.
- Backend engineers building batch inference scripts who want 429 cooldown and automatic key switching.
- Platform engineers setting up automation pipelines who need primary, daily, and fallback quota tiers.
- LLM gateway maintainers who need to inspect key-pool status and validate provider configuration.
Related Skills
Provides Claw with character-library selection, switching, saving, and global SOUL.md style sync for role-based conversation.
An AIONE Agentic AI Infrastructure SDK wrapper for building production AI agents with memory, skills, workflows, and hooks.
A Python/TypeScript SDK wrapper for the DeepSeek-Reasonix native AI coding agent, with prefix-cache support.
A browser automation tool for analysts, operators, and developers that locates elements, fills forms, extracts structured content, and supports no-code scheduling and export.