Compare LLM Models
Select 2-5 models to compare side by side.
Select models to compare
0 / 5 selected
Select 2-5 models
| Attribute | |
|---|---|
| Company | |
| Release Date | 2026-04-21 |
| Parameters (B) | 27 |
| Architecture | Hybrid Gated DeltaNet-Attention Transformer with Vision Encoder |
| Context Window | 262144 |
| Input Modalities | text, image, video |
| Open Source | Yes |
| License | apache-2.0 |
| Score | 8.7 |
| VRAM (GB) | 28 |
| Compute | Official FP8 weights (~25 GB) need roughly 28–32 GB VRAM for stable single-GPU inference; recommended setups include RTX 5090 (32 GB), H100, or dual 24 GB GPUs (e.g., 2× RTX 4090) with tensor parallelism. Best served via vLLM or SGLang with MTP speculative decoding on FP8-capable NVIDIA GPUs. |
| Benchmark: GSM8K | 93.3 |
| Benchmark: HumanEval | 83.9 |
| Benchmark: MMLU | 86.2 |
| Pricing (Input) | — |
| Pricing (Output) | — |