Mercury 2.5 Preview
PreviewThe latest diffusion large language model (dLLM) from Inception, released Aug 31, 2026. Instead of generating tokens sequentially, Mercury 2.5 produces and refines many tokens in parallel, reaching ~1,107 tokens/sec on standard GPUs. Inception positions it as the fastest reasoning LLM, reporting a 10+ point intelligence gain over Mercury 2 and quality comparable to cost-optimized frontier models (GPT-5.6 Luna Low, Gemini 3.5 Flash-Lite, Claude Haiku 4.5). It supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output, and targets latency-sensitive production workloads such as search agents, voice pipelines, and coding subagents. 260K-token context, up to 65,536 output tokens; API-only (proprietary) via Inception and OpenRouter (inception/mercury-2.5-preview). List pricing is $0.20/$0.75 per Mtok input/output (a launch promo ran at $0.04/$0.15). Architecture recorded as unknown: it is a diffusion LM rather than a standard dense/MoE autoregressive transformer.
Specifications
- License
- Proprietary
- Weights
- Not released
- Architecture
- unknown
- Parameters
- Undisclosed
- Context window
- 260K tokens
- Max output
- 66K tokens
- Knowledge cutoff
- β
- Price (in / out, $/M)
- $0.2 / $0.75
- Modalities
- TextCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.