AI Agent Hub
Back to models
Mistral: Mistral Small 4 (batch) logo

Mistral: Mistral Small 4 (batch)

Open Source mistralai Released 2026-03-16
11.0 / 100 119.0B params 256K context Proprietary

About this model

Mistral Small 4 (batch) is the asynchronous batch inference offering for Mistral Small 4 (API id mistral-small-2603), released on March 16, 2026. It runs the same 119B-parameter mixture-of-experts model (6.5B active parameters per token, 128 experts with 4 routed) as the standard chat API, including configurable reasoning effort, native function calling, and text-plus-image inputs with text outputs. Batch jobs trade latency for lower per-token cost and are suited to large offline evaluation, enrichment, and document-processing pipelines.

The underlying model unifies instruct chat, Magistral-style reasoning, Devstral-style coding agents, and Pixtral-style vision in one Apache 2.0 open-weight checkpoint (mistralai/Mistral-Small-4-119B-2603). A 256k-token context window supports long documents and multi-turn agents; reasoning_effort can be set to none for fast responses or high for harder math, code, and analysis tasks.

On public benchmarks reported at launch, Mistral Small 4 with high reasoning reaches strong scores on MMLU-Pro, GPQA-Diamond, LiveCodeBench, AIME 2025, and MMMU-Pro while emitting shorter completions than many comparably sized open models on long-context and coding tasks. Self-hosters typically serve the weights with vLLM using multi-GPU tensor parallelism; the batch product itself is accessed only through Mistral's Batch API.

Benchmark Scores

MMLU-Pro
78.0
MMMU-Pro
60.0
AIME-2025
83.8
GPQA-Diamond
71.2
LiveCodeBench
63.6

Technical Specs

  • Parameters: 119.0B
  • Architecture: Mixture-of-Experts (MoE)
  • Context Window: 256,000 tokens
  • Input Modalities: text, image

Hardware Requirements

  • Compute: API only

Pricing

Input Output Currency
0.07 / 1M tokens 0.30 / 1M tokens USD