AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
Company
Release Date 2024-09-19
Parameters (B) 1.5
Architecture Dense decoder-only Transformer (RoPE, SwiGLU, RMSNorm, GQA)
Context Window 32768
Input Modalities text
Open Source Yes
License apache-2.0
Score 7.3
VRAM (GB) 3
Compute Single GPU with ~3 GB VRAM for BF16 inference (official Qwen benchmark: 2.95 GB at 1-token input); ~1.2 GB with INT4/GPTQ or AWQ. Runs comfortably on consumer GPUs (RTX 3060 12GB, RTX 4090), Apple Silicon via MLX, or CPU through Ollama/llama.cpp.
Benchmark: GSM8K 73.2
Benchmark: HumanEval 61.6
Benchmark: MMLU 60.1
Pricing (Input)
Pricing (Output)