Mistral: Mistral Small 4 (batch)
About this model
Mistral Small 4 (batch) is the asynchronous batch inference offering for Mistral Small 4 (API id mistral-small-2603), released on March 16, 2026. It runs the same 119B-parameter mixture-of-experts model (6.5B active parameters per token, 128 experts with 4 routed) as the standard chat API, including configurable reasoning effort, native function calling, and text-plus-image inputs with text outputs. Batch jobs trade latency for lower per-token cost and are suited to large offline evaluation, enrichment, and document-processing pipelines.
The underlying model unifies instruct chat, Magistral-style reasoning, Devstral-style coding agents, and Pixtral-style vision in one Apache 2.0 open-weight checkpoint (mistralai/Mistral-Small-4-119B-2603). A 256k-token context window supports long documents and multi-turn agents; reasoning_effort can be set to none for fast responses or high for harder math, code, and analysis tasks.
On public benchmarks reported at launch, Mistral Small 4 with high reasoning reaches strong scores on MMLU-Pro, GPQA-Diamond, LiveCodeBench, AIME 2025, and MMMU-Pro while emitting shorter completions than many comparably sized open models on long-context and coding tasks. Self-hosters typically serve the weights with vLLM using multi-GPU tensor parallelism; the batch product itself is accessed only through Mistral's Batch API.
Benchmark Scores
Technical Specs
- Parameters: 119.0B
- Architecture: Mixture-of-Experts (MoE)
- Context Window: 256,000 tokens
- Input Modalities: text, image
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.07 / 1M tokens | 0.30 / 1M tokens | USD |