AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
Company
Release Date 2026-04-21
Parameters (B) 27
Architecture Hybrid Gated DeltaNet-Attention Transformer with Vision Encoder
Context Window 262144
Input Modalities text, image, video
Open Source Yes
License apache-2.0
Score 8.7
VRAM (GB) 28
Compute Official FP8 weights (~25 GB) need roughly 28–32 GB VRAM for stable single-GPU inference; recommended setups include RTX 5090 (32 GB), H100, or dual 24 GB GPUs (e.g., 2× RTX 4090) with tensor parallelism. Best served via vLLM or SGLang with MTP speculative decoding on FP8-capable NVIDIA GPUs.
Benchmark: GSM8K 93.3
Benchmark: HumanEval 83.9
Benchmark: MMLU 86.2
Pricing (Input)
Pricing (Output)