Xiaomi: MiMo-V2.5
0.0 / 10
1.1M context
Proprietary
About this model
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
Technical Specs
- Architecture: text+image+audio+video->text
- Context Window: 1,050,000 tokens
- Input Modalities: text
Hardware Requirements
- API-only (no local hardware needed)
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.14 / 1M tokens | 0.28 / 1M tokens | USD |