AI Agent Hub
Back to models
🤖

Xiaomi: MiMo-V2.6-Pro-UltraSpeed

Multimodal xiaomi Released 2026-09-21
46.0 / 100 1020.0B params 1M context Proprietary

About this model

MiMo-V2.6-Pro-UltraSpeed is Xiaomi's latency-optimized serving mode for the MiMo-V2.6-Pro flagship. It uses the same sparse MoE checkpoint (about 1.02 trillion total parameters with roughly 42 billion active per token) as the standard Pro API and the open MiMo-V2.6-Pro-RL weights on Hugging Face, but routes inference through Xiaomi's TileRT stack with aggressive throughput optimizations. Xiaomi positions it for real-time coding agents, interactive assistants, and other production workflows where wall-clock latency matters, claiming up to roughly 20x faster output than standard MiMo-V2.6-Pro while preserving flagship-level quality on complex tasks.

The underlying V2.6 series is built as a native omni-modal, long-horizon model with up to 1M-token context, supporting text, image, video, and audio inputs with text output, deep thinking, tool calling, structured outputs, and streaming. MiMo-V2.6-Pro-RL extends the family with large-scale mixed reinforcement learning across coding, general agents, visual, and cybersecurity domains. Vendor-reported agent evaluations for MiMo-V2.6 Pro include strong scores on Terminal Bench 2.1, Toolathlon-Verified, DeepSWE, AutomationBench, and CyberGym, reflecting a focus on autonomous software engineering and security workflows.

UltraSpeed is offered primarily through the Xiaomi MiMo API platform (and resellers such as OpenRouter) rather than as a separate downloadable weight variant; self-hosters typically deploy MiMo-V2.6-Pro-RL with cluster-scale SGLang or vLLM configurations. Pricing is materially higher than standard Pro per token, reflecting the premium for sustained high throughput on Xiaomi's managed infrastructure.

Benchmark Scores

HLE
49.4
DeepSWE
71.9
CyberGym
94.0
GDPval-AA
1673.0
AutomationBench
53.1
Agents-Last-Exam
31.6
Terminal-Bench-2.1
89.9
Toolathlon-Verified
76.9

Technical Specs

  • Parameters: 1020.0B
  • Architecture: Sparse MoE Transformer
  • Context Window: 1,048,576 tokens
  • Input Modalities: text, image, audio

Hardware Requirements

  • Compute: API only

Pricing

Input Output Currency
4.35 / 1M tokens 8.70 / 1M tokens USD