Model family timeline
Last updated Sep 14, 2026
DeepSeek V4 model releases
A source-backed timeline for the DeepSeek V4 model family, collecting release dates, labs, access details, context windows, and major lifecycle changes.
6 models
DeepSeek-V4.1-Flash
AvailableDeepSeek's efficient, very-low-cost flagship released September 10 2026, retiring V4-Flash and taking over V4-Pro API traffic on September 14 (DeepSeek reports it beats V4-Pro on performance, cost, speed, and total time). A 552B-total-parameter multimodal sparse Mixture-of-Experts model built on a new Causal Encoder-Decoder architecture that activates only ~8B parameters per token on input and ~16B on output for cheaper long-context prefill, with a 1M-token context, up to 384K output tokens, and native image understanding (vision in, text out). The other headline change over July's V4-Flash is memory: FP4 quantization plus "pure CSA2" cross-layer attention reuse compress the KV cache to roughly 890 bytes per token — about an 8x reduction, cutting HBM to a quarter and SSD to an eighth for an equivalent conversation state, which is what makes million-token agentic runs practical on a single node. Open-weight under the MIT license, downloadable and self-hostable, and served across many inference providers. API pricing is $0.30/$1.20 per Mtok input/output at peak (01:00-04:00 and 06:00-10:00 UTC weekdays) and half that off-peak ($0.15/$0.60), with cache hits around $0.006/$0.003 per Mtok. DeepSeek reports it narrowly edges Claude Opus 5 and GPT-5.6 Sol on DeepSWE, but there is no independent benchmark table at launch and vendor performance claims are unverified.
DeepSeek-V4-Flash-Vision-Exp
RetiredDeepSeek's first multimodal V4 model — an experimental vision-understanding checkpoint that went live on the DeepSeek API (model='deepseek-v4-flash-vision-exp') on Aug 21, 2026. It extends DeepSeek-V4-Flash with image understanding while keeping its full text capabilities (agents, reasoning, coding, and world knowledge), matching V4-Flash on text benchmarks. DeepSeek reports a major jump on multimodal agent benchmarks over V4-Flash, bringing multimodal-agent performance close to Opus-4.8 — a vendor-reported result, unverified by an independent harness at launch. Keeps V4-Flash's 284B-total / 13B-active sparse MoE architecture and 1M-token context, with up to ~393K output tokens; accepts text plus up to 600 images per request (8,192px per side, 64 MiB payload) and returns text only, with images billed at up to 384 tokens each. API pricing held at V4-Flash rates: $0.22 / $0.66 per 1M input/output tokens, with a $0.007 per 1M cached-input rate. API-only and experimental at launch — weights were not published, so treated as proprietary. RETIRED 2026-09-10: superseded by DeepSeek-V4.1-Flash, whose native multimodal support absorbs this experiment. For compatibility the `deepseek-v4-flash-vision-exp` API id temporarily routes to V4.1-Flash.
DeepSeek-V4-Pro-0813
DeprecatedThe dated GA build behind DeepSeek's 'deepseek-v4-pro' API id, pinned on the official pricing table with an OpenRouter listing dated Aug 12 2026 — the Pro-tier counterpart to the way V4-Flash graduated as DeepSeek-V4-Flash-0731. It is a quiet version pin (no separate launch post or benchmark card), keeping the V4-Pro architecture and the 1M-token context / ~384K max-output window. List pricing is cache-heavy: $0.435 input cache-miss / $0.003625 cache-hit / $0.87 output per Mtok, with concurrency 500 (vs Flash's 2500); DeepSeek warns a significant, undated API price increase is coming. Thinking is on by default at effort 'high' (requested medium/xhigh both collapse to high). Architecture figures (1.6T total / 49B active MoE, hybrid long-context attention) are carried over from the April 2026 V4-Pro preview and are not independently reconfirmed for the 0813 build; Hugging Face still hosts only the April preview weights (MIT), with no confirmed separate 0813 open-weight repo, so this row is recorded as proprietary/API-only. DEPRECATED 2026-09-14: the `deepseek-v4-pro` API id this dated build served now routes to DeepSeek-V4.1-Flash at Flash rates, pending the V4.1-Pro launch.
DeepSeek-V4-Flash-0731
RetiredThe production release of DeepSeek's V4-Flash tier — the April V4-Flash preview retrained on a substantially improved post-training pipeline targeting coding, agents, reasoning, and tool use, with no change to the base architecture. Retains 284B total / 13B active parameters (MoE) and the 1M-token context window. DeepSeek reports the 0731 build scoring higher than its own larger V4-Pro-Preview on all nine agent and coding benchmarks it published — a vendor-reported result, with independent replication still limited at launch. Weights released on Hugging Face under the MIT license; API pricing held at $0.14 / $0.28 per Mtok. The upgrade is silent for existing callers: same endpoint, same key, same deepseek-v4-flash model name, zero migration cost. RETIRED 2026-09-10: the `deepseek-v4-flash` API id this dated build served was retired in favour of DeepSeek-V4.1-Flash and now routes there.
DeepSeek V4-Flash
RetiredEfficient V4 companion model with 284B total / 13B active parameters and the same one-million-token context window. RETIRED 2026-09-10: superseded by DeepSeek-V4.1-Flash. For compatibility the `deepseek-v4-flash` API id temporarily routes to V4.1-Flash.
DeepSeek V4-Pro
DeprecatedPreview-series sparse MoE flagship with a one-million-token context window and 1.6T total / 49B active parameters. DEPRECATED 2026-09-14: from 04:00 UTC all `deepseek-v4-pro` requests route to DeepSeek-V4.1-Flash, billed at V4.1-Flash rates, and will continue to until V4.1-Pro launches. The id still answers, so this is a redirect rather than a retirement.