AI Agent Hub
Back to models
Inception: Mercury 2.5 logo

Inception: Mercury 2.5

Closed Source inception Released 2026-09-08
12.0 / 100 260K context Proprietary

About this model

Mercury 2.5 is Inception Labs' flagship production diffusion large language model (dLLM). Instead of classic left-to-right autoregressive decoding, it generates and refines many tokens in parallel through iterative denoising, which Inception pairs with tunable reasoning effort, parallel tool calls, and schema-aligned JSON for agentic and structured workflows. Inception positions it as a major quality step over Mercury 2 while preserving a low-latency, cost-efficient serving profile, with a 260K-token context window and reported throughput up to roughly 1,100 output tokens per second on widely available NVIDIA GPUs.

The model targets high-volume, latency-sensitive production paths such as search and RAG pipelines, voice agents, coding subagents, customer support, and enterprise search. Weights are not publicly released; access is via the Inception API and third-party gateways including Baseten and OpenRouter, with list pricing of $0.20 per million input tokens and $0.75 per million output tokens (often discounted at launch). Inception describes Mercury 2.5 as comparable in quality tier to cost-optimized frontier models while remaining optimized for workloads where many small model calls compound latency and cost.

Reported evaluation highlights from Inception's launch materials and independent runs include strong GPQA-Diamond performance and meaningful agentic capability on long-context and tool-use style tasks, while third-party suites such as Artificial Analysis report more modest Humanity's Last Exam scores relative to larger frontier models. Mercury 2.5 is text-in, text-out only and is marketed primarily as a reasoning model for developers building agents and interactive applications rather than as an open-weight foundation model.

Benchmark Scores

HLE
11.8
GPQA-Diamond
79.0
Terminal-Bench-2.1
34.5

Technical Specs

  • Architecture: Diffusion language model (dLLM)
  • Context Window: 260,000 tokens
  • Input Modalities: text

Hardware Requirements

  • Compute: API only

Pricing

Input Output Currency
0.04 / 1M tokens 0.15 / 1M tokens USD