Mercury 2.5 reaches 1,107 tokens per second
Inception released Mercury 2.5, a diffusion language model positioned for production systems where many model calls can compound latency. The company says the model is the largest diffusion LLM ever trained and reports performance of 1,107 tokens per second on widely available NVIDIA GPUs, with a 260K token context window.
Mercury 2.5 is framed as a cheap-and-fast alternative to small frontier models rather than a replacement for flagship reasoning systems. Inception claims a 40% intelligence gain over Mercury 2 and lists support for tunable reasoning, parallel tool calls, and schema-aligned JSON output. Pricing is $0.20 per million input tokens and $0.75 per million output tokens, with an 80% launch discount dropping those to $0.04 input / $0.15 output.
The model is aimed at search, RAG, voice agents, tool-use loops, context compaction, query rewriting, and routing. OpenCall reported that Mercury reduced P99 response time from several minutes to one second, while Augment Code said context compaction latency fell 82%, from roughly 150 seconds to 27 seconds, with cost down 90% and no quality regression.
Mercury 2.5 is available through the Inception API, OpenRouter, and Baseten, using OpenAI-compatible endpoints. Inception also previewed Mercury Voice for voice agents with time-to-first-token under 170 milliseconds and Mercury Router for routing prompts across open and closed models.