AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
Company
Release Date 2026-05-28
Parameters (B) 35
Architecture Mixture-of-Experts (MoE) with Hybrid Attention
Context Window 262144
Input Modalities text, image, video
Open Source Yes
License apache-2.0
Score 8.5
VRAM (GB) 24
Compute NVIDIA Hopper (H100/H200) or Blackwell (B200/GB200/GB300/DGX Spark GB10) GPU with ≥24 GB VRAM for practical inference; Linux; vLLM ≥0.28.0 with NVIDIA ModelOpt NVFP4 (--quantization modelopt); FlashInfer attention and Marlin MoE backend recommended
Benchmark: GSM8K 96.3
Benchmark: HumanEval 40.6
Benchmark: MMLU 85
Pricing (Input)
Pricing (Output)