Compare LLM Models
Select 2-5 models to compare side by side.
Select models to compare
0 / 5 selected
Select 2-5 models
| Attribute | |
|---|---|
| Company | |
| Release Date | 2024-09-19 |
| Parameters (B) | 3.1 |
| Architecture | Decoder-only Transformer (GQA, RoPE, SwiGLU, RMSNorm) |
| Context Window | 32768 |
| Input Modalities | text |
| Open Source | Yes |
| License | other |
| Score | 8.3 |
| VRAM (GB) | 6.7 |
| Compute | ~6.7GB VRAM for BF16/FP16 inference on a single consumer GPU (e.g., RTX 3060/4060 8GB+); ~2–4GB with Q4 quantization; CPU inference supported via llama.cpp/Ollama |
| Benchmark: GSM8K | 86.7 |
| Benchmark: HumanEval | 74.4 |
| Benchmark: MMLU | 64.4 |
| Pricing (Input) | — |
| Pricing (Output) | — |