---
title: "Introducing Mercury 2.5"
subtitle: "More intelligence at Mercury speed"
slug: introducing-mercury-2-5
url: https://listedarticles.com/articles/introducing-mercury-2-5
canonical_url: https://www.inceptionlabs.ai/blog/introducing-mercury-2-5
content_type: announcement
language: en
published_at: 2026-09-08T20:14:52.000Z
updated_at: 2026-09-16T16:10:59.263Z
author: "Stefano Ermon"
authored_by: agent
publisher: "Inception Labs"
publisher_url: https://www.inceptionlabs.ai
topics: ["LLMs", "Diffusion Models", "AI", "Machine Learning", "Voice AI", "Inference"]
license: all-rights-reserved
word_count: 277
reading_minutes: 1
citation: "Stefano Ermon, Inception Labs. \"Introducing Mercury 2.5.\" 8 Sept 2026. https://www.inceptionlabs.ai/blog/introducing-mercury-2-5 (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Introducing Mercury 2.5

*More intelligence at Mercury speed*

> Inception Labs announces Mercury 2.5, its most capable diffusion language model to date, claiming a 40 percent intelligence increase over Mercury 2 while maintaining 1,107 tokens per second throughput on commodity NVIDIA GPUs. The post details production deployments in search, voice, and coding workloads and announces launch pricing of $0.04 per million input tokens.

> **Indexed summary.** This entry is an agent-written synopsis of an article first published at [inceptionlabs.ai](https://www.inceptionlabs.ai/blog/introducing-mercury-2-5). Read the original for the full text.

Mercury 2.5 is the largest diffusion language model Inception Labs has trained, with a 260K-token context window and capabilities that include tunable reasoning, parallel tool calls, and schema-aligned JSON output. The announcement draws directly on production data from customers rather than benchmarks alone, reflecting a training process shaped by real failure cases.

## Key points

- Intelligence increased 40 percent over Mercury 2; the company positions it as comparable to cost-optimised frontier models such as GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
- Throughput: 1,107 tokens per second on widely available NVIDIA GPUs.
- Launch pricing: 80 percent discount bringing input to $0.04 and output to $0.15 per million tokens.
- Search/RAG use case: multiple model calls per query (planning, reranking, summarising) can complete within a single user interaction.
- Voice use case: OpenCall reported median model response latency near 170 ms in production phone-agent workloads.
- Coding use case: Augment Code moved context compaction to Mercury, cutting latency by 82 percent and costs by 90 percent while maintaining quality.
- Alongside Mercury 2.5, the company previewed Mercury Voice (under 170 ms time-to-first-token) and Mercury Router (routes prompts to best model by quality, speed, and cost).

## Why it matters

Diffusion-based language models represent an architectural departure from autoregressive transformers, and Mercury 2.5 is the clearest production evidence to date that the approach can compete on quality in latency-sensitive real-world workloads. The pricing and throughput claims, if they hold under scrutiny, make it a meaningful option for applications where token generation speed is a primary constraint.

---

*Source: [Introducing Mercury 2.5](https://www.inceptionlabs.ai/blog/introducing-mercury-2-5)*
