Meta: Llama 3.3 70B Instruct
About this model
Llama 3.3 70B Instruct is Meta's December 2024 instruction-tuned release in the Llama 3 family. It is a dense, text-only autoregressive model with about 70 billion parameters, Grouped-Query Attention (GQA), and a 128K-token context window. Weights were pretrained on roughly 15 trillion tokens from public sources (knowledge cutoff December 2023), then aligned with supervised fine-tuning, RLHF, and large-scale synthetic instruction data for multilingual dialogue, coding, math, and tool use.
On Meta's published English benchmarks, the model matches or beats Llama 3.1 70B Instruct on most tasks and approaches Llama 3.1 405B Instruct on several metrics despite far lower compute at inference. Reported instruction-tuned scores include 86.0 on MMLU (0-shot CoT), 68.9 on MMLU-Pro, 92.1 on IFEval, 50.5 on GPQA Diamond, 88.4 on HumanEval, 87.6 on MBPP EvalPlus, and 77.0 on MATH. It supports eight languages (English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai) and multiple tool-calling formats via standard chat templates.
The checkpoint is released under the Llama 3.3 Community License for research and commercial use (subject to license terms and acceptable-use policy). It is intended for assistant-style chat and can be deployed with Hugging Face Transformers, vLLM, or Meta's reference stack, with optional 8-bit or 4-bit loading for memory-constrained GPUs. Developers should pair the model with application-level safety guardrails such as Llama Guard where appropriate.
Benchmark Scores
Technical Specs
- Parameters: 70.0B
- Architecture: Transformer (GQA)
- Context Window: 128,000 tokens
- Input Modalities: text
Hardware Requirements
- VRAM: 80.0 GB
- Compute: 2x NVIDIA A100 80GB (BF16) or 1x A100 80GB (8-bit)
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.22 / 1M tokens | 0.50 / 1M tokens | USD |