Meta: Llama 3.2 3B Instruct
About this model
Llama 3.2 3B Instruct is Meta’s instruction-tuned, text-only member of the Llama 3.2 family, with about 3.21 billion parameters and a 128K-token context window. It uses an optimized autoregressive Transformer with grouped-query attention (GQA), trained on a multilingual mix of publicly available data (knowledge cutoff December 2023) and aligned with supervised fine-tuning, rejection sampling, and direct preference optimization for helpful, safe dialogue. Official support covers English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai, with broader language exposure during pretraining.
The model targets on-device and resource-constrained deployment: agentic retrieval, summarization, rewriting, and lightweight tool use without the footprint of 8B+ models. On Meta’s published English benchmarks, it reaches 63.4% on MMLU (5-shot), 77.7% on GSM8K (CoT), 48.0% on MATH (CoT), 78.6% on ARC-Challenge, 77.4% on IFEval, and 32.8% on GPQA, with strong tool-use scores on BFCL v2 (67.0%). Quantized SpinQuant and QLoRA variants are also released for mobile and CPU inference via ExecuTorch.
Released under the Llama 3.2 Community License on September 25, 2024, Llama 3.2 3B Instruct is intended for commercial and research chat and agent applications. Deployers should pair it with appropriate safety guardrails (e.g., Llama Guard, Prompt Guard) and run task-specific evaluations, especially when extending beyond the eight supported languages or deploying in high-risk settings.
Benchmark Scores
Technical Specs
- Parameters: 3.21B
- Architecture: Transformer
- Context Window: 128,000 tokens
- Input Modalities: text
Hardware Requirements
- VRAM: 8.0 GB
- Compute: Single NVIDIA GPU with 8GB VRAM
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.05 / 1M tokens | 0.33 / 1M tokens | USD |