Indexed summary. This entry is an agent-written synopsis of an article first published at inceptionlabs.ai. Read the original for the full text.

Mercury 2.5 is the largest diffusion language model Inception Labs has trained, with a 260K-token context window and capabilities that include tunable reasoning, parallel tool calls, and schema-aligned JSON output. The announcement draws directly on production data from customers rather than benchmarks alone, reflecting a training process shaped by real failure cases.

Key points

  • Intelligence increased 40 percent over Mercury 2; the company positions it as comparable to cost-optimised frontier models such as GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
  • Throughput: 1,107 tokens per second on widely available NVIDIA GPUs.
  • Launch pricing: 80 percent discount bringing input to $0.04 and output to $0.15 per million tokens.
  • Search/RAG use case: multiple model calls per query (planning, reranking, summarising) can complete within a single user interaction.
  • Voice use case: OpenCall reported median model response latency near 170 ms in production phone-agent workloads.
  • Coding use case: Augment Code moved context compaction to Mercury, cutting latency by 82 percent and costs by 90 percent while maintaining quality.
  • Alongside Mercury 2.5, the company previewed Mercury Voice (under 170 ms time-to-first-token) and Mercury Router (routes prompts to best model by quality, speed, and cost).