Facts
- Announced
- Mercury 2.5 · 2026-09-09
- Noted
- a diffusion-based language model · 2026-09-09
- Noted
- refines multiple tokens in parallel · 2026-09-09
- Noted
- outputs 1107 tokens per second · 2026-09-09
- Noted
- up from Mercury 2's 1009 · 2026-09-09
- Noted
- context length doubled to 260,000 tokens · 2026-09-09
- Noted
- rates its quality on par with GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5 · 2026-09-09
- Noted
- Output pricing stays at $0.75 per million tokens · 2026-09-09
- Noted
- input pricing drops to $0.20 · 2026-09-09
- Noted
- 80 percent launch discount on OpenRouter · 2026-09-09
Structured graph also available as JSON at /public/entities/mercury-2-5.
CC BY 4.0.
12h ago
Inception announced Mercury 2.5, a diffusion-based language model that refines multiple tokens in parallel instead of generating text left to right. It outputs 1107 tokens per second, up from Mercury 2's 1009, with context length doubled to 260,000 tokens. Inception rates its quality on par with GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Output pricing stays at $0.75 per million tokens while input pricing drops to $0.20, with an 80 percent launch discount on OpenRouter.