Inception Ships Mercury 2.5 Diffusion LLM at 1107 Tokens Per Second
Mercury 2.5 puts diffusion-based generation on par with the cost-efficient frontier models GPT-5.6 Luna Low and Gemini 3.5 Flash-Lite while cutting input pricing, making parallel token generation a practical option for latency-sensitive agent workloads.
Reporting from 1 source: GIGAZINE.
Inception announced Mercury 2.5, a diffusion-based language model that refines multiple tokens in parallel instead of generating text left to right. It outputs 1107 tokens per second, up from Mercury 2's 1009, with context length doubled to 260,000 tokens. Inception rates its quality on par with GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Output pricing stays at $0.75 per million tokens while input pricing drops to $0.20, with an 80 percent launch discount on OpenRouter.
Mercury 2.5 applies the diffusion approach used in image generation to text, building a rough draft of the whole answer and refining multiple tokens in parallel instead of deciding one token at a time. Inception says the model matches the quality of GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
The company cites two deployments. AI phone agent developer OpenCall cut median response time to about 170 milliseconds, and coding AI firm Augment Code shortened context compression from about 150 seconds to 27 seconds with a 90 percent cost reduction.
Inception also previewed Mercury Voice for speech AI and Mercury Router, which sends input to the appropriate model. The company says training on its largest next-generation model has begun, with a release targeted within the next few months.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.