Mistral: Mistral Large 3 2512 (batch)
About this model
Mistral Large 3 2512 (batch) is the asynchronous batch inference offering for Mistral’s flagship open-weight model mistral-large-2512. It exposes the same Mistral Large 3 675B Instruct checkpoint used in real-time APIs: a granular sparse mixture-of-experts multimodal language model with about 41B active parameters per token, 675B total parameters, a 256K-token context window, native vision understanding, function calling, structured outputs, and broad multilingual support under the Apache 2.0 license.
The batch endpoint is intended for high-volume, latency-tolerant workloads such as offline evaluation, dataset labeling, bulk summarization, and large-scale content generation. Requests are queued and processed at reduced cost relative to synchronous chat completions while preserving the same model behavior, safety settings, and tool-use capabilities. Typical use cases mirror the standard model: enterprise assistants, retrieval-augmented generation over long documents, agentic workflows, and multimodal document Q&A.
Open weights can be self-hosted on modern NVIDIA clusters (for example FP8 on 8x H200 or NVFP4 on 8x H100/A100), but the batch catalog entry targets Mistral’s managed Batch API rather than on-premise deployment. Performance on knowledge, reasoning, coding, and agent benchmarks aligns with publicly reported Mistral Large 3 2512 results from the Hugging Face model card, Mistral documentation, Artificial Analysis, Vals AI, and DataLearner leaderboards.
Benchmark Scores
Technical Specs
- Parameters: 675.0B
- Architecture: Sparse Mixture-of-Experts
- Context Window: 256,000 tokens
- Input Modalities: text, image
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.25 / 1M tokens | 0.75 / 1M tokens | USD |