AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
Company
Release Date 2024-09-19
Parameters (B) 3.1
Architecture Decoder-only Transformer (GQA, RoPE, SwiGLU, RMSNorm)
Context Window 32768
Input Modalities text
Open Source Yes
License other
Score 8.3
VRAM (GB) 6.7
Compute ~6.7GB VRAM for BF16/FP16 inference on a single consumer GPU (e.g., RTX 3060/4060 8GB+); ~2–4GB with Q4 quantization; CPU inference supported via llama.cpp/Ollama
Benchmark: GSM8K 86.7
Benchmark: HumanEval 74.4
Benchmark: MMLU 64.4
Pricing (Input)
Pricing (Output)