Inception released Mercury 2.5, its latest diffusion-based language model. It is listed on OpenRouter from 8 September 2026. The model writes several words at once instead of one after another. Inception claims 1,107 tokens per second on standard GPUs. It reports a gain of more than 10 points in quality over Mercury 2. Inception places the quality near cost-optimised frontier models from OpenAI, Google and Anthropic. The model takes 260,000 tokens of input and returns up to 65,536. It supports adjustable reasoning effort, parallel tool calls and JSON output matching a schema. One provider hosts it, currently at a discounted price. Inception aims it at search agents, voice pipelines and coding subagents.
What changed
Mercury 2 scored more than 10 points lower on intelligence.
What it unlocks
Running latency-sensitive agents and voice pipelines on a very fast, low-cost reasoning model.
- $0.04 / $0.15 per 1M tokens
- 1,107 tokens/sec claimed
- 260K token context
- 80% off list price
What you need to act on it
- an API key via OpenRouter or Inception
Sources