AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
Company
Release Date 2026-04-02
Parameters (B) 30.7
Architecture Dense Multimodal Transformer
Context Window 256000
Input Modalities text, image
Open Source Yes
License apache-2.0
Score 8.5
VRAM (GB) 70
Compute Full BF16 weights require ~70GB VRAM (single 80GB NVIDIA H100). Q4_0 quantization reduces requirement to ~17.5GB (24GB consumer GPU); Q8/SFP8 needs ~35GB. Supports Ollama, vLLM, and Transformers on Linux/macOS.
Benchmark: GSM8K 85
Benchmark: HumanEval 85
Benchmark: MMLU 85.2
Pricing (Input)
Pricing (Output)