AI Agent Hub
Back to models
🤖

Xiaomi: MiMo-V2.6-Pro

Open Source xiaomi Released 2026-09-22
46.0 / 100 1020.0B params 1M context Proprietary

About this model

MiMo-V2.6-Pro is Xiaomi MiMo’s flagship open-weight omni-modal reasoning model, released in September 2026 under the MIT license as the MiMo-V2.6-Pro-RL checkpoint on Hugging Face and ModelScope. It is built as a sparse mixture-of-experts stack with about 1.02 trillion total parameters and roughly 42 billion activated parameters per token, combining a 70-layer language backbone (sliding-window and global attention), dedicated vision and audio encoders, and a multi-token prediction speculative decoder. Inputs span text, images, video, and audio with up to 1M tokens of context, targeting long-horizon agent workflows in software engineering, desktop automation, cybersecurity, and multimodal tool use.

Training emphasizes scaled reinforcement learning rather than separate per-domain RL runs: a single mixed GRPO stage across coding, general agents, visual agents, and security tasks, with groupwise agentic grading and aligned RL safeguards. Official evaluations highlight strong agentic performance—DeepSWE v1.1 at 71.9, Terminal-Bench 2.1 at 89.9, Toolathlon-Verified at 76.9, AutomationBench at 53.1, and CyberGym at 94.0—alongside GDPval-AA 2.1 at 1673 and Agents’ Last Exam at 31.6, placing it near frontier closed models on many harnesses while leading open-source systems on the Artificial Analysis composite index at launch.

Weights are intended for large-cluster serving (official SGLang and vLLM recipes use multi-GPU tensor and expert parallelism) and are also available through the Xiaomi MiMo Open Platform API, MiMo Studio, and OpenRouter. The series documents large-scale RL post-training (on the order of 750k trajectories for Pro) as a step toward recursive self-improvement on verifiable, environment-grounded tasks.

Benchmark Scores

DeepSWE
71.9
CyberGym
94.0
GDPval-AA
1673.0
AutomationBench
53.1
Agents-Last-Exam
31.6
Terminal-Bench-2.1
89.9
Toolathlon-Verified
76.9

Technical Specs

  • Parameters: 1020.0B
  • Architecture: Sparse MoE Transformer
  • Context Window: 1,000,000 tokens
  • Input Modalities: text, image, audio

Hardware Requirements

  • VRAM: 1280.0 GB
  • Compute: 2-node cluster with 16× NVIDIA H100 80GB (tensor parallel 16, expert parallel 16)

Pricing

Input Output Currency
0.43 / 1M tokens 0.87 / 1M tokens USD