Meta: Llama Guard 4 12B
About this model
Llama Guard 4 12B is Meta's natively multimodal content safety classifier, released alongside the Llama 4 family. It has 12 billion parameters in a dense early-fusion Transformer architecture pruned from the Llama 4 Scout pre-trained checkpoint by removing routed experts and router layers while keeping the shared expert feedforward blocks, then fine-tuned for safety classification without additional pre-training. The model classifies both LLM prompts and responses as safe or unsafe and, when unsafe, names violated hazard categories aligned with the MLCommons hazards taxonomy (S1-S14, including code interpreter abuse). It unifies English and multilingual text moderation from Llama Guard 3-8B with single- and multi-image understanding from Llama Guard 3-11B-vision, supporting up to several images per prompt and integration with the Llama Moderations API.
Operationally, Llama Guard 4 acts as a conditional generation model: it consumes chat-formatted text and optional images and emits short textual safety judgments. Meta reports that on in-house safety test sets for output filtering, it matches or exceeds Llama Guard 3 on English (69% recall, 11% false positive rate, 61% F1), multilingual text, single-image, and multi-image settings, with the largest gains on multi-image prompts. It is intended for input filtering, output filtering, or both in production pipelines guarding generative models such as Llama 4 Scout and Maverick.
The checkpoint is gated on Hugging Face under the Llama 4 Community License. Inference is designed for a single GPU at bfloat16 with roughly 24 GB VRAM and a 163,840-token context window, sharing tokenizer and vision encoder components with Llama 4 Scout and Maverick. Limitations include dependence on training data for policy coverage, reduced reliability on categories needing fresh factual knowledge, and susceptibility to adversarial or prompt-injection attacks; Meta recommends complementary defenses such as Llama Prompt Guard 2 where appropriate.
Technical Specs
- Parameters: 12.0B
- Architecture: Dense early-fusion Transformer
- Context Window: 163,840 tokens
- Input Modalities: text, image
Hardware Requirements
- VRAM: 24.0 GB
- Compute: Single NVIDIA GPU with 24GB VRAM
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.18 / 1M tokens | 0.18 / 1M tokens | USD |