Meta: Llama 3.2 1B Instruct
About this model
Llama 3.2 1B Instruct is Meta's smallest instruction-tuned model in the Llama 3.2 text lineup, with about 1.23 billion parameters and a 128,000-token context window. It is built on an optimized Transformer architecture with grouped-query attention (GQA), trained on a multilingual mixture of publicly available data (up to roughly 9 trillion tokens, knowledge cutoff December 2023), and aligned with supervised fine-tuning, rejection sampling, and direct preference optimization similar to Llama 3.1. The model targets efficient on-device and edge deployment: dialogue, retrieval-augmented generation, summarization, query rewriting, and lightweight agentic workflows in English and seven additional supported languages (German, French, Italian, Portuguese, Hindi, Spanish, and Thai).
Meta positions this checkpoint for commercial and research use under the Llama 3.2 Community License. Official evaluations on the bf16 instruct variant report solid gains over the 1B pretrained base on instruction following and reasoning-oriented tasks—for example MMLU (49.3%), IFEval (59.5%), GSM8K chain-of-thought (44.4%), and MATH chain-of-thought (30.6%)—while remaining compact enough for low-latency inference. The family also ships heavily quantized variants (4-bit weights with SpinQuant/QLoRA) aimed at mobile CPUs and constrained hardware, making this model a common choice when latency, memory, and cost matter more than peak frontier capability.
Compared with larger Llama 3.2 and 3.1 models, the 1B instruct model trades depth of knowledge and complex reasoning for speed and accessibility. It is best suited to high-volume chat, classification, short-form generation, and orchestration layers that call tools or retrieve context, rather than demanding multi-step math proofs or expert-level science QA. Developers typically run it in half precision or quantized form on a single modest GPU or on-device runtime stacks such as ExecuTorch.
Benchmark Scores
Technical Specs
- Parameters: 1.23B
- Architecture: Transformer (GQA)
- Context Window: 128,000 tokens
- Input Modalities: text
Hardware Requirements
- VRAM: 4.0 GB
- Compute: Single NVIDIA GPU with 4GB VRAM (8GB recommended for 128k context)
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.03 / 1M tokens | 0.20 / 1M tokens | USD |