AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
Company
Release Date 2026-04-02
Parameters (B) 25.2
Architecture Mixture-of-Experts (MoE) Transformer
Context Window 256000
Input Modalities text, image
Open Source Yes
License apache-2.0
Score 8.4
VRAM (GB) 57.7
Compute BF16 inference requires ~58 GB VRAM (e.g., A100 80GB); 8-bit quantization ~29 GB (RTX 4090/A6000); Q4_0 quantization ~14.4 GB on consumer GPUs. MoE architecture loads all 25.2B parameters but activates only 3.8B per token.
Benchmark: GSM8K 68
Benchmark: HumanEval 76.8
Benchmark: MMLU 82.6
Pricing (Input)
Pricing (Output)