AI Agent Hub
Back to models
Meta: Llama Guard 4 12B logo

Meta: Llama Guard 4 12B

Multimodal meta-llama Released 2025-04-29
-- 12.0B params 163.8K context Proprietary

About this model

Llama Guard 4 12B is Meta's natively multimodal content safety classifier, released alongside the Llama 4 family. It has 12 billion parameters in a dense early-fusion Transformer architecture pruned from the Llama 4 Scout pre-trained checkpoint by removing routed experts and router layers while keeping the shared expert feedforward blocks, then fine-tuned for safety classification without additional pre-training. The model classifies both LLM prompts and responses as safe or unsafe and, when unsafe, names violated hazard categories aligned with the MLCommons hazards taxonomy (S1-S14, including code interpreter abuse). It unifies English and multilingual text moderation from Llama Guard 3-8B with single- and multi-image understanding from Llama Guard 3-11B-vision, supporting up to several images per prompt and integration with the Llama Moderations API.

Operationally, Llama Guard 4 acts as a conditional generation model: it consumes chat-formatted text and optional images and emits short textual safety judgments. Meta reports that on in-house safety test sets for output filtering, it matches or exceeds Llama Guard 3 on English (69% recall, 11% false positive rate, 61% F1), multilingual text, single-image, and multi-image settings, with the largest gains on multi-image prompts. It is intended for input filtering, output filtering, or both in production pipelines guarding generative models such as Llama 4 Scout and Maverick.

The checkpoint is gated on Hugging Face under the Llama 4 Community License. Inference is designed for a single GPU at bfloat16 with roughly 24 GB VRAM and a 163,840-token context window, sharing tokenizer and vision encoder components with Llama 4 Scout and Maverick. Limitations include dependence on training data for policy coverage, reduced reliability on categories needing fresh factual knowledge, and susceptibility to adversarial or prompt-injection attacks; Meta recommends complementary defenses such as Llama Prompt Guard 2 where appropriate.

Technical Specs

  • Parameters: 12.0B
  • Architecture: Dense early-fusion Transformer
  • Context Window: 163,840 tokens
  • Input Modalities: text, image

Hardware Requirements

  • VRAM: 24.0 GB
  • Compute: Single NVIDIA GPU with 24GB VRAM

Pricing

Input Output Currency
0.18 / 1M tokens 0.18 / 1M tokens USD