AI Agent Hub
Back to models
🤖

Xiaomi: MiMo-V2.6-Flash

Multimodal xiaomi Released 2026-09-22
38.0 / 100 309.0B params 1M context Proprietary

About this model

MiMo-V2.6-Flash is Xiaomi MiMo's efficiency-oriented member of the MiMo-V2.6 family, released in September 2026 under the MIT license with open weights on Hugging Face (MiMo-V2.6-Flash-RL). It is a native omnimodal sparse mixture-of-experts model with 309 billion total parameters and about 15 billion activated per token, combining text, image, video, and audio understanding with text generation and a one-million-token context window. The series emphasizes large-scale mixed reinforcement learning across coding agents, general tool use, visual tasks, and cybersecurity, using asynchronous GRPO, groupwise agentic grading, and optional MOPD distillation for hard-to-verify domains.

On Xiaomi's published agent and security evaluations, Flash delivers strong terminal and workflow scores— including 87.6 on Terminal-Bench 2.1, 73.6 on Toolathlon-Verified, 67.9 on DeepSWE v1.1, 52.3 on AutomationBench, and a leading 95.1 on CyberGym—while staying cheaper to run than MiMo-V2.6-Pro. Independent Artificial Analysis testing reports 35.1% on Humanity's Last Exam and 73.1% on MMMU-Pro. The model supports deep thinking, tool calling, structured output, streaming, and web search, and is served via the Xiaomi MiMo API (OpenAI-compatible), MiMo Studio, MiMo Desktop, and OpenRouter, with SGLang and vLLM recipes for self-hosting.

Architecturally, Flash uses a 48-layer hybrid backbone (sliding-window and global attention) with 256 routed experts (8 active), dedicated MiMo ViT and audio encoders, and a five-layer multi-token prediction speculative decoder for higher throughput. It targets high-frequency API calls, long-horizon agent harnesses, and multimodal professional workflows where Pro-level capability is not required at Pro-level cost.

Benchmark Scores

HLE
35.1
DeepSWE
67.9
CyberGym
95.1
MMMU-Pro
73.1
AutomationBench
52.3
Agents-Last-Exam
27.6
Terminal-Bench-2.1
87.6
Toolathlon-Verified
73.6

Technical Specs

  • Parameters: 309.0B
  • Architecture: Sparse MoE
  • Context Window: 1,000,000 tokens
  • Input Modalities: text, image, audio

Hardware Requirements

  • VRAM: 640.0 GB
  • Compute: 8x NVIDIA H100 80GB (SGLang TP8 recommended)

Pricing

Input Output Currency
0.14 / 1M tokens 0.28 / 1M tokens USD