AI Agent Hub
Back to models
Meta: Llama 3.3 70B Instruct logo

Meta: Llama 3.3 70B Instruct

Open Source meta-llama Released 2024-12-06
8.0 / 100 70.0B params 128K context Proprietary

About this model

Llama 3.3 70B Instruct is Meta's December 2024 instruction-tuned release in the Llama 3 family. It is a dense, text-only autoregressive model with about 70 billion parameters, Grouped-Query Attention (GQA), and a 128K-token context window. Weights were pretrained on roughly 15 trillion tokens from public sources (knowledge cutoff December 2023), then aligned with supervised fine-tuning, RLHF, and large-scale synthetic instruction data for multilingual dialogue, coding, math, and tool use.

On Meta's published English benchmarks, the model matches or beats Llama 3.1 70B Instruct on most tasks and approaches Llama 3.1 405B Instruct on several metrics despite far lower compute at inference. Reported instruction-tuned scores include 86.0 on MMLU (0-shot CoT), 68.9 on MMLU-Pro, 92.1 on IFEval, 50.5 on GPQA Diamond, 88.4 on HumanEval, 87.6 on MBPP EvalPlus, and 77.0 on MATH. It supports eight languages (English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai) and multiple tool-calling formats via standard chat templates.

The checkpoint is released under the Llama 3.3 Community License for research and commercial use (subject to license terms and acceptable-use policy). It is intended for assistant-style chat and can be deployed with Hugging Face Transformers, vLLM, or Meta's reference stack, with optional 8-bit or 4-bit loading for memory-constrained GPUs. Developers should pair the model with application-level safety guardrails such as Llama Guard where appropriate.

Benchmark Scores

MATH
77.0
MBPP
87.6
MMLU
86.0
IFEval
92.1
MMLU-Pro
68.9
HumanEval
88.4
GPQA-Diamond
50.5

Technical Specs

  • Parameters: 70.0B
  • Architecture: Transformer (GQA)
  • Context Window: 128,000 tokens
  • Input Modalities: text

Hardware Requirements

  • VRAM: 80.0 GB
  • Compute: 2x NVIDIA A100 80GB (BF16) or 1x A100 80GB (8-bit)

Pricing

Input Output Currency
0.22 / 1M tokens 0.50 / 1M tokens USD