AI Agent Hub
Back to models
Mistral: Ministral 3 8B 2512 (batch) logo

Mistral: Ministral 3 8B 2512 (batch)

Multimodal mistralai Released 2025-12-02
5.0 / 100 8.0B params 262.1K context Proprietary

About this model

Ministral 3 8B 2512 (batch) is Mistral AI’s batched API offering for the Ministral 3 8B instruct checkpoint released in December 2025. It is a dense, Apache 2.0 open-weight multimodal model built via cascade distillation from Mistral Small 3.1, combining an ~8.4B language backbone with a ~0.4B vision encoder for text-and-image understanding. The family targets efficient edge and local deployment while supporting a very long context window, native tool calling, structured JSON output, and broad multilingual coverage.

The batch endpoint exposes the same capabilities as the standard Ministral 8B instruct API—vision, agentic tool use, and chat/instruction following—at reduced per-token cost for asynchronous workloads. Weights are published on Hugging Face in FP8 and BF16 variants (alongside separate Base and Reasoning checkpoints), and the model is designed to run on a single consumer or datacenter GPU when quantized appropriately.

On public evaluations, Ministral 3 8B delivers strong results for its size class on knowledge and math benchmarks, competitive multimodal scores, and solid code metrics on the dedicated Reasoning variant, while API-side leaderboards (e.g., Artificial Analysis) report instruct-style scores on MMLU-Pro, GPQA, HLE, and agentic suites. It is well suited to local assistants, document and image QA, translation, and lightweight agent pipelines where cost and latency matter.

Benchmark Scores

ARC
88.0
HLE
4.3
MATH
87.6
MBPP
70.0
MMLU
76.1
MMMU
55.1
AGIEval
59.1
MMLU-Pro
64.2
MMMU-Pro
46.0
AIME-2024
86.0
AIME-2025
78.7
GPQA-Diamond
66.8
LiveCodeBench
61.6
Terminal-Bench-2.1
4.1

Technical Specs

  • Parameters: 8.0B
  • Architecture: Transformer
  • Context Window: 262,144 tokens
  • Input Modalities: text, image

Hardware Requirements

  • VRAM: 24.0 GB
  • Compute: Single NVIDIA GPU with 24GB VRAM (BF16); 12GB VRAM with FP8 weights

Pricing

Input Output Currency
0.07 / 1M tokens 0.07 / 1M tokens USD