Xiaomi: MiMo-V2.6-Pro-UltraSpeed
About this model
MiMo-V2.6-Pro-UltraSpeed is Xiaomi's latency-optimized serving mode for the MiMo-V2.6-Pro flagship. It uses the same sparse MoE checkpoint (about 1.02 trillion total parameters with roughly 42 billion active per token) as the standard Pro API and the open MiMo-V2.6-Pro-RL weights on Hugging Face, but routes inference through Xiaomi's TileRT stack with aggressive throughput optimizations. Xiaomi positions it for real-time coding agents, interactive assistants, and other production workflows where wall-clock latency matters, claiming up to roughly 20x faster output than standard MiMo-V2.6-Pro while preserving flagship-level quality on complex tasks.
The underlying V2.6 series is built as a native omni-modal, long-horizon model with up to 1M-token context, supporting text, image, video, and audio inputs with text output, deep thinking, tool calling, structured outputs, and streaming. MiMo-V2.6-Pro-RL extends the family with large-scale mixed reinforcement learning across coding, general agents, visual, and cybersecurity domains. Vendor-reported agent evaluations for MiMo-V2.6 Pro include strong scores on Terminal Bench 2.1, Toolathlon-Verified, DeepSWE, AutomationBench, and CyberGym, reflecting a focus on autonomous software engineering and security workflows.
UltraSpeed is offered primarily through the Xiaomi MiMo API platform (and resellers such as OpenRouter) rather than as a separate downloadable weight variant; self-hosters typically deploy MiMo-V2.6-Pro-RL with cluster-scale SGLang or vLLM configurations. Pricing is materially higher than standard Pro per token, reflecting the premium for sustained high throughput on Xiaomi's managed infrastructure.
Benchmark Scores
Technical Specs
- Parameters: 1020.0B
- Architecture: Sparse MoE Transformer
- Context Window: 1,048,576 tokens
- Input Modalities: text, image, audio
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| 4.35 / 1M tokens | 8.70 / 1M tokens | USD |