DeepSeek-V4-Flash-Vision-Exp
PreviewDeepSeek's first multimodal V4 model — an experimental vision-understanding checkpoint that went live on the DeepSeek API (model='deepseek-v4-flash-vision-exp') on Aug 21, 2026. It extends DeepSeek-V4-Flash with image understanding while keeping its full text capabilities (agents, reasoning, coding, and world knowledge), matching V4-Flash on text benchmarks. DeepSeek reports a major jump on multimodal agent benchmarks over V4-Flash, bringing multimodal-agent performance close to Opus-4.8 — a vendor-reported result, unverified by an independent harness at launch. Keeps V4-Flash's 284B-total / 13B-active sparse MoE architecture and 1M-token context, with up to ~393K output tokens; accepts text plus up to 600 images per request (8,192px per side, 64 MiB payload) and returns text only, with images billed at up to 384 tokens each. API pricing held at V4-Flash rates: $0.22 / $0.66 per 1M input/output tokens, with a $0.007 per 1M cached-input rate. API-only and experimental at launch — weights were not published, so treated as proprietary.
Specifications
- License
- Proprietary
- Weights
- Not released
- Architecture
- Mixture-of-Experts
- Parameters
- 284B · 13B active
- Context window
- 1M tokens
- Max output
- 393K tokens
- Knowledge cutoff
- —
- Price (in / out, $/M)
- $0.22 / $0.66
- Modalities
- TextVisionCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.