LLM Releases

Lab release history

Last updated Sep 2, 2026

Alibaba (Qwen) model releases

Qwen team at Alibaba Cloud; prolific open-weight releases. This page collects the lab's model releases, lifecycle events, source links, and model metadata in one crawlable record.

29
Models
1
Labs
23
Open
4
Recent

29 models

Qwen3.8-Max-0902

Available
Alibaba (Qwen)FrontierProprietary

An updated snapshot of Alibaba's flagship Qwen3.8-Max, released Sep 2 2026 (model id qwen3.8-max-0902; alias qwen3.8-max-2026-09-02). A post-training upgrade on the unchanged 2.4-trillion-parameter Mixture-of-Experts base (~95B active per token), it sharpens coding for engineering-scale projects and long-horizon autonomous development, multi-tool agent orchestration, and native vision (chart reasoning, document parsing, multimodal perception). Accepts text, image, and video input and returns text, with a 1M-token context, up to ~131K output tokens, and an optional thinking mode carrying up to 256K chain-of-thought tokens. Pricing is unchanged from Qwen3.8-Max at $2.00 / $6.00 per 1M input/output tokens. Vendor-reported gains include a +22-point CodeArena jump to 1,691 (first on that leaderboard at launch); figures are unverified. Served API-first (Alibaba, OpenRouter); the -Max flagship tier is closed-weight, so treat this snapshot as proprietary despite the open-weight base family. Distinct from the base Qwen3.8-Max (2026-08-03).

MoE2.4T1M ctxSep 2, 2026

Qwen3.8-Flash

Available
Alibaba (Qwen)Proprietary

The productionized, managed Qwen Cloud API model announced Aug 26 2026 alongside the open-weight Qwen3.8-Flash-Next, with official QwenCloud pricing confirmed Aug 27-28. It runs the same Qwen4-preview architecture as Flash-Next — a 125B-total sparse Mixture-of-Experts that activates ~6B parameters per token (~180B stored once a 51B n-gram embedding table and a multi-token-prediction module are counted; roughly 95% sparsity), built on Gated-DeltaNet + Qwen Sparse Attention with a gated residual stream and Muon-trained large linear layers — but ships as the hosted service rather than the self-host weights. On QwenCloud it defaults to a 1M-token context with built-in tools and accepts text/image/video in, returning text out. List pricing is $0.15 input / $0.47 output / $0.016 cache-hit per Mtok (domestic China Y0.8/Y2.7/Y0.1), roughly a third of DeepSeek-V4-Flash and about one-thirteenth of the Qwen3.8-Max flagship. Also reachable through OpenCode Go's flat-rate subscription. Distinct catalog row from the open-weight Qwen3.8-Flash-Next (self-host, Qwen Community License 1.0); this managed API is recorded proprietary/API-only. Self-reported benchmarks carry over from Flash-Next (SWE-bench Pro 62.5, DeepSWE 58.7) — all vendor numbers, unverified by independent labs at launch. Thinking on by default.

MoE125B1M ctxAug 26, 2026

Qwen3.8-Flash-Next

Preview
Alibaba (Qwen)Open weights

An open-weight, experimental preview of the architecture that will underpin Qwen4, released Aug 26 2026 (Qwen/Qwen3.8-Flash-Next). A sparse MoE with ~6B active parameters (headline 125B-with-6B-activated; ~180B stored once a 51B n-gram embedding table and 4B multi-token-prediction module are counted), 512 experts (10 routed + 1 shared), and a hybrid Gated-DeltaNet + Qwen Sparse Attention design. Native 262,144-token context, extensible to 1M via YaRN. Accepts text, image, and video in and returns text out. Distinct from the managed Qwen Cloud 'Qwen3.8-Flash' API (which defaults to 1M context and bundled tools); this Next build is catalog/self-host only with no hosted list price at launch, served via Transformers, vLLM, SGLang, and TokenSpeed. Weights under the Qwen Community License 1.0. Self-reported vs DeepSeek-V4-Flash-0731: DeepSWE 58.7 vs 54.4, SWE-bench Pro 62.5 vs 56.0, LiveCodeBench v6 91.9, GPQA Diamond 91.7, though NL2Repo 48.1 vs 54.2 is a regression; vision self-reports include AndroidWorld 84.5 and RealWorldQA 88.5 — all vendor numbers, unverified at launch. Thinking on by default.

MoE125B262K ctxAug 26, 2026

Qwen3.8-27B

Available
Alibaba (Qwen)Open weights

The open-weight, single-GPU sibling of Qwen3.8-Max, published by Alibaba on Hugging Face on Aug 14 2026 under Apache 2.0 — the smaller open release Alibaba had promised alongside the closed Qwen3.8-Max flagship. A 27B dense model (~28B counting the ~1B vision encoder) with 64 layers, hidden size 5,120, and a 248,320-token vocabulary. Uses a hybrid attention stack — 48 Gated DeltaNet linear-attention layers to 16 full Gated Attention layers (a 3:1 split) — for a native 262,144-token context, extendable to 1M via YaRN. Natively multimodal (text, image, and video input; text output) and ships with Multi-Token Prediction for speculative decoding. Quantized (Unsloth dynamic GGUFs) it runs in ~16-17GB of VRAM, fitting a single consumer GPU such as a 3090 or 4090 — positioned as one of the most capable local models of 2026.

Hybrid27B262K ctxAug 14, 2026

Qwen3.8-Max

Available
Alibaba (Qwen)FrontierOpen weights

Alibaba's largest model to date and the flagship of the Qwen3.8 line — a 2.4-trillion-parameter sparse Mixture-of-Experts (~95B active per query) that Alibaba positions just behind Anthropic's Fable 5 on overall performance. Previewed 2026-07-19 at the World AI Conference in Shanghai, it went generally available on 2026-08-03 with a published benchmark table, standard API access, and firm per-token pricing ($2 / $6 per 1M input/output tokens). Fully multimodal (text, image, video input) over a 1M-token context. Alibaba also committed to shipping open weights for both Qwen3.8-Max and a smaller Qwen3.8-27B checkpoint.

MoE2.4T1.0M ctxAug 3, 2026

Qwen3.7-Flash

Available
Alibaba (Qwen)Proprietary

Cost-optimized multimodal member of the Qwen3.7 line — a vision-language reasoning model with a 1M-token context, tuned for high-volume multimodal agent workloads (visual coding, screen perception, browser/computer use, search) where cost matters more than peak intelligence. Launched quietly on Jul 27, 2026 as an OpenRouter/API listing at $0.03/$0.13 per 1M tokens, making it the cheapest 1M-context multimodal model available at release. Closed-weights and API-only; Alibaba published no technical report, benchmark suite, or architecture details, though community speculation points to a small sparse-MoE design.

MoEUndisc.1M ctxJul 27, 2026

Qwen3.7-Plus

Available
Alibaba (Qwen)Proprietary

Multimodal sibling of Qwen3.7-Max that adds vision input and GUI grounding for screen perception, browser automation, and hybrid GUI+CLI agent workflows. 1M-token context; closed-weights and API-only. Previewed at the May 2026 Alibaba Cloud Summit and reached general availability in June 2026 at a low price point ($0.40/$1.60 per 1M tokens).

MoEUndisc.1M ctxJun 3, 2026

Qwen3.7-Max

Available
Alibaba (Qwen)FrontierProprietary

Alibaba's proprietary flagship in the Qwen3.7 "Agent Frontier" line — a text-only sparse-MoE model with a 1M-token context, tuned for long-horizon agentic, coding, and reasoning workloads. Parameter count is undisclosed; access is API-only via Alibaba Cloud Model Studio / DashScope (and aggregators such as OpenRouter).

MoEUndisc.1M ctxMay 20, 2026

Qwen3.6-27B

Available
Alibaba (Qwen)Open source

Dense 27B that punches far above its weight on agentic coding — easy to self-host on a single GPU node.

Dense27B256K ctxMay 12, 2026

Qwen3.5-9B

Available
Alibaba (Qwen)Open source

The flagship of Alibaba's small dense Qwen3.5 models. Independent analysis (Artificial Analysis) rated it the most intelligent model under 10B parameters at launch — roughly double the score of the next-closest sub-10B models — and the most intelligent multimodal model under 15B, leading peers on MMMU-Pro (~69%). A dense 9B with native vision, a 262K-token context, and the Qwen3.5 family's unified hybrid thinking / non-thinking mode. Native weights are BF16; in 4-bit it needs ~6GB, within reach of consumer laptops. High intelligence comes with heavy reasoning token usage (~260M output tokens to run the Intelligence Index).

Dense9B262K ctxMar 2, 2026

Qwen3.5-4B

Available
Alibaba (Qwen)Open source

A dense 4B in Alibaba's small Qwen3.5 family, rated by Artificial Analysis as the most intelligent model under 5B parameters at launch — outscoring several 7B–9B peers despite roughly half the parameters. Native vision, a 262K-token context, and the family's hybrid thinking / non-thinking mode; Apache-2.0 licensed. Scores ~65% on MMMU-Pro multimodal reasoning and runs in ~3GB at 4-bit, suitable for lightweight on-device agents.

Dense4B262K ctxMar 2, 2026

Qwen3.5-2B

Available
Alibaba (Qwen)Open source

A dense 2B Qwen3.5 model built for high-throughput, low-latency edge and on-device use. Despite its size it matches a 7B-class peer on Artificial Analysis's Intelligence Index. Apache-2.0, with native vision, a 262K-token context, and the family's hybrid thinking / non-thinking mode; runs in under 2GB at 4-bit, fitting laptops and smartphones.

Dense2B262K ctxMar 2, 2026

Qwen3.5-0.8B

Available
Alibaba (Qwen)Open source

The smallest Qwen3.5 model — a dense 0.8B designed for the most constrained on-device deployments, operating in non-thinking (instruct) mode by default. Apache-2.0, with native vision, a 262K-token context, and the family's hybrid thinking / non-thinking mode; needs roughly 2GB of VRAM and runs under 2GB at 4-bit, targeting smartphones and embedded hardware. Notable for a sub-1B model, it still scores ~26% on MMMU-Pro multimodal reasoning.

Dense0.8B262K ctxMar 2, 2026

Qwen3.5-397B

Available
Alibaba (Qwen)FrontierOpen source

Native vision-language MoE supporting 201 languages with a 1M-token context.

MoE397B1M ctxFeb 20, 2026

Qwen3-Coder-Next

Available
Alibaba (Qwen)Open source

Apache-licensed Qwen3-Next coding-agent model with 80B total / 3B active parameters, 256K context, and long-horizon tool-use training.

Hybrid80B262K ctxFeb 3, 2026

Qwen3-Coder-480B-A35B-Instruct

Available
Alibaba (Qwen)Open source

Alibaba Qwen's large open coding-agent model: a 480B-total / 35B-active MoE released under Apache-2.0, tuned for code generation, repository-level software engineering, tool calling, and long-horizon agent workflows with a 256K-token native context.

MoE480B262K ctxJul 22, 2025

Qwen3-235B-A22B

Available
Alibaba (Qwen)Open source

Largest open Qwen3 MoE, introducing hybrid thinking/non-thinking modes and 119-language coverage.

MoE235B128K ctxApr 28, 2025

Qwen2.5-Omni-7B

Available
Alibaba (Qwen)Open weights

Local omni-modal Qwen model that supports text, image, audio, video, and speech generation in a 7B package.

Dense7B ctxMar 26, 2025

Qwen2.5-Max

Available
Alibaba (Qwen)Proprietary

Proprietary MoE flagship for the Qwen2.5 generation, released through Qwen Chat and Alibaba Cloud APIs.

MoEUndisc. ctxJan 29, 2025

Qwen2.5-VL-72B

Available
Alibaba (Qwen)Open weights

Vision-language Qwen2.5 model for image, document, video, and agentic visual grounding tasks.

Dense72B128K ctxJan 26, 2025

QwQ-32B-Preview

Available
Alibaba (Qwen)Open source

Qwen's first public reasoning-preview model, aimed at math, coding, and deliberate problem solving.

Dense32B32K ctxNov 28, 2024

Qwen2.5-Coder-32B

Available
Alibaba (Qwen)Open source

Code-specialized Qwen2.5 model family, with the 32B checkpoint as the flagship open coding model.

Dense32B128K ctxNov 12, 2024

Qwen2.5-72B

Available
Alibaba (Qwen)Open weights

Broad Qwen2.5 foundation-model update spanning general, coding, math, and multimodal descendants.

Dense72B128K ctxSep 19, 2024

Qwen2-72B

Available
Alibaba (Qwen)Open weights

Qwen2's largest dense model, introducing stronger multilingual support, coding/math gains, and long-context variants.

Dense72B128K ctxJun 7, 2024

Qwen1.5-110B

Available
Alibaba (Qwen)Open weights

Largest Qwen1.5 model, released as the bridge from the original Qwen line to Qwen2.

Dense110B32K ctxFeb 5, 2024

Qwen1.5-72B-Chat

Available
Alibaba (Qwen)Open weights

Largest chat-tuned Qwen1.5 dense checkpoint, released with stronger human-preference alignment, multilingual support, and 32K context.

Dense72B33K ctxFeb 4, 2024

Qwen-72B

Available
Alibaba (Qwen)Open weights

Alibaba's first major open Qwen model and the start of a prolific open-weight line.

Dense72B32K ctxNov 30, 2023

Qwen-14B

Available
Alibaba (Qwen)Open weights

Second open Qwen size, expanding the first-generation Qwen language-model lineup.

Dense14B8K ctxSep 25, 2023

Qwen-7B

Available
Alibaba (Qwen)Open weights

Alibaba's first open Qwen checkpoint and the start of the Qwen open-model line.

Dense7B32K ctxAug 3, 2023