AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
Company
Release Date 2025-07-31
Parameters (B) 30.5
Architecture Mixture-of-Experts Transformer
Context Window 262144
Input Modalities text
Open Source Yes
License apache-2.0
Score 8.4
VRAM (GB) 20.4
Compute Single GPU with ~20 GB VRAM for Q4_K_M GGUF (e.g. RTX 3090/4090); MoE architecture activates only 3.3B of 30.5B params per token for efficient inference; 24 GB+ recommended for comfortable KV-cache headroom at moderate context lengths
Benchmark: GSM8K 90.7
Benchmark: HumanEval 92.7
Benchmark: MMLU 84.7
Pricing (Input)
Pricing (Output)