AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
Company
Release Date 2024-09-19
Parameters (B) 7.6
Architecture Transformer
Context Window 131072
Input Modalities text
Open Source Yes
License apache-2.0
Score 8.5
VRAM (GB) 16
Compute ~16 GB VRAM for FP16/BF16 inference on a single consumer GPU (e.g., RTX 3090/4080); ~6 GB with Q4 quantization; CPU inference possible via llama.cpp/Ollama with sufficient system RAM
Benchmark: GSM8K 91.6
Benchmark: HumanEval 84.8
Benchmark: MMLU 75.4
Pricing (Input)
Pricing (Output)