LLM Releases
← Catalog

Mercury 2.5 Preview

Preview
InceptionProprietary

The latest diffusion large language model (dLLM) from Inception, released Aug 31, 2026. Instead of generating tokens sequentially, Mercury 2.5 produces and refines many tokens in parallel, reaching ~1,107 tokens/sec on standard GPUs. Inception positions it as the fastest reasoning LLM, reporting a 10+ point intelligence gain over Mercury 2 and quality comparable to cost-optimized frontier models (GPT-5.6 Luna Low, Gemini 3.5 Flash-Lite, Claude Haiku 4.5). It supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output, and targets latency-sensitive production workloads such as search agents, voice pipelines, and coding subagents. 260K-token context, up to 65,536 output tokens; API-only (proprietary) via Inception and OpenRouter (inception/mercury-2.5-preview). List pricing is $0.20/$0.75 per Mtok input/output (a launch promo ran at $0.04/$0.15). Architecture recorded as unknown: it is a diffusion LM rather than a standard dense/MoE autoregressive transformer.

Specifications

License
Proprietary
Weights
Not released
Architecture
unknown
Parameters
Undisclosed
Context window
260K tokens
Max output
66K tokens
Knowledge cutoff
β€”
Price (in / out, $/M)
$0.2 / $0.75
Modalities
TextCode

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.