Compare LLM Models
Select 2-5 models to compare side by side.
Select models to compare
0 / 5 selected
Select 2-5 models
| Attribute | |
|---|---|
| Company | |
| Release Date | 2025-04-29 |
| Parameters (B) | 0.6 |
| Architecture | Dense Transformer (GQA) |
| Context Window | 32768 |
| Input Modalities | text |
| Open Source | Yes |
| License | apache-2.0 |
| Score | 7.2 |
| VRAM (GB) | 2 |
| Compute | Runs on CPU or any modern GPU with ≥2 GB VRAM at FP16 (≈1.3 GB weights at 4K context); Q4_K_M quantization needs ~0.5 GB weights and fits integrated or low-end GPUs. Full 32K context adds ~3.8 GB KV cache (~6 GB total). |
| Benchmark: GSM8K | 59.6 |
| Benchmark: HumanEval | 31 |
| Benchmark: MMLU | 52.8 |
| Pricing (Input) | — |
| Pricing (Output) | — |