NVDA 222.27 ▲1.34%GOOGL 349.54 ▲0.64%MSFT 493.78 ▼0.80%AMD 559.82 ▲2.70%INTC 108.60 ▼0.18%TSMC 434.67 ▲1.03%AMZN 253.71 ▲1.00%META 665.75 ▼2.43%AAPL 336.13 ▼0.26%PLTR 177.64 ▲0.79%
Markets at last close

Models

Mercury 2.5 reaches 1,107 tokens per second

·1 min read

Inception released Mercury 2.5, a diffusion language model positioned for production systems where many model calls can compound latency. The company says the model is the largest diffusion LLM ever trained and reports performance of 1,107 tokens per second on widely available NVIDIA GPUs, with a 260K token context window.

Mercury 2.5 is framed as a cheap-and-fast alternative to small frontier models rather than a replacement for flagship reasoning systems. Inception claims a 40% intelligence gain over Mercury 2 and lists support for tunable reasoning, parallel tool calls, and schema-aligned JSON output. Pricing is $0.20 per million input tokens and $0.75 per million output tokens, with an 80% launch discount dropping those to $0.04 input / $0.15 output.

The model is aimed at search, RAG, voice agents, tool-use loops, context compaction, query rewriting, and routing. OpenCall reported that Mercury reduced P99 response time from several minutes to one second, while Augment Code said context compaction latency fell 82%, from roughly 150 seconds to 27 seconds, with cost down 90% and no quality regression.

Mercury 2.5 is available through the Inception API, OpenRouter, and Baseten, using OpenAI-compatible endpoints. Inception also previewed Mercury Voice for voice agents with time-to-first-token under 170 milliseconds and Mercury Router for routing prompts across open and closed models.

Originally reported by alphasignal.aiRead the source →
Related coverage