Compare LLM Models
Select 2-5 models to compare side by side.
Select models to compare
0 / 5 selected
Select 2-5 models
| Attribute | |
|---|---|
| Company | |
| Release Date | 2026-05-28 |
| Parameters (B) | 35 |
| Architecture | Mixture-of-Experts (MoE) with Hybrid Attention |
| Context Window | 262144 |
| Input Modalities | text, image, video |
| Open Source | Yes |
| License | apache-2.0 |
| Score | 8.5 |
| VRAM (GB) | 24 |
| Compute | NVIDIA Hopper (H100/H200) or Blackwell (B200/GB200/GB300/DGX Spark GB10) GPU with ≥24 GB VRAM for practical inference; Linux; vLLM ≥0.28.0 with NVIDIA ModelOpt NVFP4 (--quantization modelopt); FlashInfer attention and Marlin MoE backend recommended |
| Benchmark: GSM8K | 96.3 |
| Benchmark: HumanEval | 40.6 |
| Benchmark: MMLU | 85 |
| Pricing (Input) | — |
| Pricing (Output) | — |