NVIDIA: Switchyard
About this model
NVIDIA NeMo Switchyard is not a single foundation model but an open-source orchestration layer for routing LLM traffic across a configurable pool of efficient and capable models. Implemented primarily in Rust with Python bindings, it exposes OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages-compatible APIs while keeping routing logic separate from provider endpoints. Built-in algorithms include passthrough, weighted random, LLM task classifiers, stage routers that react to tool outcomes, escalation flows, composite strategies, and optional learned prefill routing, so agents can balance quality, latency, and cost without hard-coding one model per step.
Published evaluations focus on routing effectiveness rather than classic knowledge benchmarks. LangChain’s Deep Agents suite (145 multi-turn agentic tasks drawn from production-style scenarios including tool use, retrieval, and long-context work) reported roughly 80.0% task completion when routing NVIDIA Nemotron 3.5 Lightning with Claude Opus 4.8 via escalation, versus 86.0% for Opus alone and 77.7% for Lightning alone, with about 7% of calls sent to the frontier model and roughly 74% lower cost than an Opus-only baseline. Cognition reported similar staged routing on FrontierCode Main at about 50.6% accuracy with materially lower mean cost than a frontier-only setup. Because the served model changes per request, fixed MMLU-style scores do not apply to Switchyard itself.
Switchyard is maintained by NVIDIA under the NeMo program (Apache-2.0 on GitHub) and is available as a self-hosted proxy or embedded library, with hosted access on OpenRouter under the identifier nvidia/switchyard. OpenRouter documents up to a 1,000,000-token effective context window bounded by whichever backend model handles a given call, with billing at the selected model’s rate rather than a separate Switchyard tariff.
Technical Specs
- Architecture: Multi-model routing orchestration
- Context Window: 1,000,000 tokens
- Input Modalities: text
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| -1000000.00 / 1M tokens | -1000000.00 / 1M tokens | USD |