AI Agent Hub
Back to skills
🤖

CN LLM Unified Router

AI Agent Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @user_1a470ba8/cn-llm-router according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

When using multiple Chinese LLM providers, engineers often deal with fragmented API keys, provider-specific request formats, uneven model capabilities, cost tracking, and concurrency limits. Manual model selection can waste budget, while batch jobs may saturate local CPU and memory. This skill consolidates 12 domestic providers behind one CLI entry point, keeping routing, invocation, billing, caching, and fallback in a single workflow.

How it works

It is a Python standard-library CLI; aside from an optional websocket-client for iFLYTEK Spark, it does not require third-party dependencies. Keys come only from environment variables and are not written to disk or packaged. The core path is prompt → task classification → strategy engine → provider adapter → model call, with side channels for cost tracking and local cache. route can produce recommendations without calling an API; chat supports streaming, system prompts, and JSON output; report aggregates spend, success rate, and latency; hardware limits concurrency and batch size based on CPU and memory; cache uses similarity plus length penalty to reduce duplicate calls. Since v2.1, auto reads model capability profiles and prefers the best match for reasoning, code, or long-context tasks; with no network or keys, a Mock mode can still exercise the pipeline.

Caveats

It is not an agent framework and does not fine-tune models, manage vector stores, or build RAG pipelines. Recommendations use a local registry and heuristic profiles, not live benchmarks. Semantic cache may misfire, so use --no-cache for finance, code, or real-time queries. Streaming token counts are often estimates; exact billing should be verified in provider consoles.

Use Cases

  • Before calling DeepSeek, Tongyi, or Kimi, use route to inspect model recommendations and cost strategy, then run chat.
  • Batch-summarize, translate, or extract long documents with streaming chat, aggregate spend in report, and use cache to avoid repeated API calls.
  • On a dev machine without network or API keys, run the chat, report, budget, and cache pipeline with --mock for debugging.
  • On a low-spec machine, rely on hardware to cap concurrency and batch size so batch jobs do not saturate CPU or memory.

Best For

  • Backend engineers maintaining keys for multiple Chinese LLM providers, who want one command to choose models and track cross-provider spend.
  • Algorithm engineers batch-processing document summaries and translations, who need cost control, concurrency limits, and offline Mock debugging.
  • Developers building agent prototypes on low-spec machines, who want safe local concurrency and network-free debugging.
  • Platform engineers managing AI budget governance, who need monthly reports, budget alerts, and troubleshooting for 429/401 errors.