Changelog
Last updated Jul 26, 2026
Everything, in order
A single feed of releases, updates, deprecations, and retractions across every lab we track. Each item links to the model and its primary source.
July 2026
- UpdatedMoonshot publishes Kimi K3 open weightsMoonshot AIsource ↗
Moonshot AI publishes the full Kimi K3 weights to Hugging Face under a Modified MIT license on July 26 — a day ahead of its announced July 27 target — making the 2.8T-parameter MoE freely downloadable, modifiable, and self-hostable, and cementing K3 as the largest open-weight model publicly available.
- ReleasedAnthropic releases Claude Opus 5Anthropicsource ↗
Anthropic releases Claude Opus 5, its flagship model for demanding reasoning, autonomous coding, and long-horizon agentic work — pitched as the go-to model for most knowledge work, approaching Fable 5 capability in many categories at about half the price. Adds a five-level 'effort' dial on the Claude API/Platform to trade compute for capability, a 1M-token context at standard pricing, and up to 128K output tokens. Standard pricing $5/$25 per Mtok (matching Opus 4.8), plus a $10/$50 Fast mode; becomes the default for Claude Max subscribers. Anthropic calls it its most aligned Opus model.
- ReleasedAnt Group's inclusionAI releases Ling-3.0-flashAnt Group (inclusionAI)source ↗
Ant Group's inclusionAI lab releases Ling-3.0-flash, a hybrid-reasoning Mixture-of-Experts model with 124B total parameters and ~5.1B active per token (1/64 expert activation), built for production-scale agents. Ant claims it matches or beats its own ~1T-parameter flagship on most benchmarks shown at 1/8 the total and 1/12 the active parameters — a vendor claim with no public benchmark table at launch. Uses a KDA + MLA hybrid-linear attention stack at a reported 5:1 ratio for an economical 256K-token context. Announced as open-weight under Apache 2.0, but weights and a model card were not yet posted to Hugging Face as of July 24; usable only via hosted API, free on OpenRouter and Vercel AI Gateway through August 3 2026.
- ReleasedGoogle releases Gemini 3.5 Flash-LiteGoogle DeepMindsource ↗
Fastest, most cost-effective 3.5-class model (~350 output tokens/s) for high-throughput agentic workloads, priced at $0.30 / $2.50 per Mtok with a 1M-token context and built-in computer use. Large step up on 3.1 Flash-Lite and beats 3 Flash on several agentic/coding evals.
- ReleasedPoolside releases Laguna S 2.1Poolsidesource ↗
Poolside releases Laguna S 2.1, a 118B-total / 8B-active open-weight MoE coding model with a 1M-token context, pitched as 'the West's most capable open-weight model' for its weight class. It scores 70.2% on Terminal-Bench 2.1 and tops the published open disclosed-size table on SWE-bench Multilingual at 78.5%. Trained in under nine weeks on 4,096 H200 GPUs, it ships weights on Hugging Face under OpenMDW-1.1 (BF16/FP8/INT4/NVFP4 + GGUF/MLX) and runs at 4-bit on a single NVIDIA DGX Spark. Hosted free at 256K context and paid at full 1M context via OpenRouter ($0.10/$0.20/$0.01 per 1M input/output/cache-read tokens).
- PreviewGoogle previews Gemini 3.5 Flash Cyber in CodeMenderGoogle DeepMindsource ↗
A cyber-specialized model built on 3.5 Flash for finding and fixing software vulnerabilities, deployed inside the CodeMender agent and reaching competitive frontier performance on CyberGym. Limited-access pilot restricted to governments and trusted partners.
- ReleasedGoogle launches Gemini 3.6 FlashGoogle DeepMindsource ↗
Google DeepMind ships Gemini 3.6 Flash, a multimodal 1M-context workhorse for agentic workflows that improves on 3.5 Flash in coding, knowledge work, and computer use while cutting output-token usage ~17% at a lower price ($1.50 / $7.50 per Mtok). Available in the Gemini API, Gemini Enterprise, and the Gemini app.
- UpdatedDeepSeek V4-Flash reaches general availabilityDeepSeeksource ↗
V4-Flash, the efficient tier of the V4 family, exits preview alongside V4-Pro as DeepSeek retires its legacy deepseek-chat and deepseek-reasoner aliases (cutoff July 24, 2026).
- UpdatedDeepSeek V4 family reaches general availabilityDeepSeeksource ↗
DeepSeek moves the V4 model family (V4-Pro and V4-Flash) out of preview and into general availability, closing a run of just under three months from the April 24 preview. The legacy deepseek-chat and deepseek-reasoner API aliases are retired on July 24, 2026 with no fallback; production traffic must reference deepseek-v4-pro / deepseek-v4-flash.
- PreviewAlibaba unveils Qwen3.8-Max-PreviewAlibaba (Qwen)source ↗
Alibaba launches Qwen3.8, a 2.4-trillion-parameter fully-multimodal model — more than double the size of its predecessor — which the Qwen team says ranks "second only to Fable 5" on overall performance (internal evals, no independent benchmarks yet). Qwen3.8-Max-Preview is available to developers via Alibaba's Token Plan subscription and the Qoder / QoderWork coding platforms, with open weights promised "soon" but no timeline, license, or architecture details disclosed.
- ReleasedMoonshot AI releases Kimi K3Moonshot AIsource ↗
Largest open-weight model to date: 2.8T-parameter MoE (896 experts, 16 active) with Kimi Delta Attention, native multimodal input, and a 1M-token context. API launched Jul 16 at $3/$15 per Mtok; full weights slated for Jul 27. Reported 93.5% GPQA Diamond (strongest published open-weight result).
- UpdatedGemini 3.5 Pro slips past its July 17 targetGoogle DeepMindsource ↗
Bloomberg reports Google delayed Gemini 3.5 Pro again after the rebuilt model fell short of internal quality goals on hallucinations and reliability — its third slipped target after June and early July. Google DeepMind's Logan Kilpatrick said on July 21 the company is testing it with partners and hopes to "land soon"; there is still no model card or public API entry.
- ReleasedThinking Machines releases InklingThinking Machines Labsource ↗
Mira Murati's lab ships its first model — a 975B-total / 41B-active multimodal MoE (text/image/audio in, text out) pretrained on ~45T tokens, released under Apache-2.0 with weights on Hugging Face (1M context) and hosted on the Tinker API (256K). Debuts at 41 on the Artificial Analysis Intelligence Index, the leading U.S. open-weights model.
- ReleasedKwaipilot ships KAT-Coder-Air V2.5Kwaipilot (Kuaishou)source ↗
The efficient ~32B-active variant of KAT-Coder V2.5, sharing the 256K context and agentic tool-use focus at roughly a fifth of Pro's price ($0.15/$0.60 per Mtok).
- ReleasedKwaipilot ships KAT-Coder-Pro V2.5Kwaipilot (Kuaishou)source ↗
Kuaishou's Kwaipilot team releases KAT-Coder-Pro V2.5, an agentic coding MoE (~72B active) trained with large-scale agentic RL in verifiable repository environments, with a 256K context, 80K max output, and API pricing of $0.74/$2.96 per Mtok via StreamLake, Atlas Cloud, and OpenRouter.
- ReleasedGPT-5.6 Luna reaches general availabilityOpenAIsource ↗
Luna, the fast, most cost-efficient tier of the GPT-5.6 family ($1/$6 per Mtok), moves from restricted preview to general availability alongside Sol and Terra.
- ReleasedMeta launches Muse Spark 1.1 — its first paid modelMeta AIsource ↗
Meta Superintelligence Labs releases Muse Spark 1.1 in US public preview on the Meta Model API, the first time Meta charges for one of its models. A natively multimodal reasoning model (text/image/video/PDF/audio in, text out) with a self-compacting 1M-token context, aimed at agentic and coding workflows and priced at $1.25/$4.25 per Mtok — about a quarter of comparable Anthropic/OpenAI models. Closed weights; benchmarks around the Opus 4.8 / GPT-5.5 tier.
- ReleasedOpenAI takes the GPT-5.6 family to GAOpenAIsource ↗
Following U.S. government review and a two-week restricted preview, OpenAI rolls out the full GPT-5.6 family — Sol, Terra, and Luna — to general availability across ChatGPT (Plus, Pro, Business, Enterprise) and the API.
- ReleasedGPT-5.6 Terra reaches general availabilityOpenAIsource ↗
Terra, the balanced everyday-work tier of the GPT-5.6 family ($2.50/$15 per Mtok), moves from restricted preview to general availability alongside Sol and Luna, available to Plus, Pro, Business, and Enterprise users in ChatGPT and via the API.
- ReleasedGPT-5.6 Sol reaches general availabilityOpenAIsource ↗
OpenAI moves the GPT-5.6 family to general availability after the June 26 restricted preview. Sol, the flagship for difficult professional, coding, research, computer-use, and tool-heavy work, ships with a 1.05M-token context window, 128K max output, and standard pricing of $5/$30 per Mtok, alongside a new max reasoning effort and ultra subagent mode.
- ReleasedSpaceXAI and Cursor launch Grok 4.5xAIsource ↗
SpaceXAI (the rebranded xAI) releases its most capable model to date — the first Grok trained jointly with Cursor (Anysphere), pitched as "Opus-class" but faster, cheaper ($2/$6 per Mtok), and ~4x more token-efficient than Opus 4.8 on SWE-bench Pro. Available in Grok Build (default), Cursor (all plans), and the SpaceXAI console; initially unavailable in the EU.
- AnnouncedMistral confirms a frontier-gap open-weight MoE with July early accessMistral AIsource ↗
CEO Arthur Mensch confirms Mistral is preparing a new open-weight Mixture-of-Experts family — "fat but sparse" — entering early access in July 2026 and aimed at the frontier open-weight tier. No parameter count, benchmarks, license terms, or release date disclosed; tracked as rumored.
- ReleasedNVIDIA releases Nemotron-Labs-3-Puzzle-75B-A9BNVIDIAsource ↗
A compressed variant of Nemotron-3-Super produced with "Iterative Puzzle", trimming the parent to 75.3B total / 9.3B active while keeping the hybrid Mamba-Transformer LatentMoE design and 1M context. NVIDIA reports ~2x higher server throughput at matched user throughput; released on Hugging Face in BF16/FP8/NVFP4 under OpenMDW-1.1.
- ReleasedTencent officially releases and open-sources Hunyuan Hy3Tencent Hunyuansource ↗
Tencent officially launches Hunyuan 3.0 (Hy3), the GA of its rebuilt third-generation model, and open-sources it under Apache-2.0. A 295B-total / 21B-active MoE (plus a 3.8B multi-token-prediction layer) with a 256K context and three selectable fast/slow inference modes; Tencent reports it rivals GLM-5.2 and DeepSeek-V4 and matches or surpasses GPT-5.5 on several science benchmarks, with 78.0 on SWE-bench Verified. Weights on Hugging Face (tencent/Hy3) and ModelScope, with a free OpenRouter route (tencent/hy3:free) through July 21, 2026.
- ReleasedPoolside releases Laguna XS 2.1Poolsidesource ↗
Poolside releases Laguna XS 2.1, an upgraded 33B-A3B open-weight MoE coding model served at 256K context. It raises SWE-bench Multilingual by 5.4 points to 63.1% over XS.2, adds open-weighted DFlash speculator models that roughly double local tokens/sec, and moves to the fully permissive OpenMDW-1.1 license. Weights are on Hugging Face (BF16/FP8/INT4/NVFP4) and it's available free on OpenRouter, with paid API pricing of $0.10/$0.20/$0.05 per 1M input/output/cache-read tokens.
- UpdatedClaude Fable 5 access restored globallyAnthropicsource ↗
Anthropic says Claude Fable 5 is available globally again on Claude Platform, Claude.ai, Claude Code, and Claude Cowork after export controls were lifted; cloud partner access is being re-enabled and safeguards now route high-risk requests to Opus 4.8.
- UpdatedClaude Mythos 5 access restored for approved partnersAnthropicsource ↗
Anthropic restored Mythos 5 access for approved U.S. organizations and continues expanding the Glasswing trusted-access program, while Mythos remains restricted rather than generally available.
June 2026
- ReleasedMeituan open-sources LongCat-2.0, trained entirely on Chinese chipsMeituan (LongCat)source ↗
Meituan releases LongCat-2.0, a 1.6T-parameter MoE (~48B active) with a 1M-token context for agentic coding, open-sourced under the MIT license on Hugging Face and GitHub. It was trained and served on a ~50,000-card cluster of domestic Chinese AI chips — the first trillion-parameter model Meituan says completed full-process training and inference on home-grown hardware. Vendor-reported results: 59.5 SWE-bench Pro (vs GPT-5.5's 58.6), 70.8 Terminal-Bench 2.1, and 77.3 SWE-bench Multilingual.
- ReleasedAnthropic releases Claude Sonnet 5Anthropicsource ↗
Anthropic releases Claude Sonnet 5, its most agentic Sonnet yet — performance approaching Opus 4.8 at a lower price, made the default model on the Free and Pro plans and available to Max, Team, and Enterprise users. Available in Claude Code and via the Claude API as claude-sonnet-5, with introductory pricing of $2 per Mtok input / $10 per Mtok output through Aug 31, 2026, then standard $3/$15.
- UpdatedWhite House lifts Anthropic Fable 5 banAnthropicsource ↗
Reporting says the White House lifted the ban on Anthropic's models after an agreement on additional safeguards, authorizing Anthropic to return Fable 5 to public release channels while Mythos remains limited to pre-vetted partners.
- ReleasedBase44 rolls out Base1, its first in-house modelBase44source ↗
Base44 (a Wix company) rolls out Base1, a general-purpose 'vibe coding' agent fine-tuned on an open-source foundation model using data from tens of millions of platform interactions, selectable alongside GPT-5.5 and Claude Opus 4.8 in its model picker.
- PreviewOpenAI previews GPT-5.6 SolOpenAIsource ↗
Sol is the highest-capability GPT-5.6 preview tier, available only to a small vetted cohort.
- UpdatedCommerce Department clears limited Claude Mythos returnAnthropicsource ↗
U.S. Commerce officials reportedly allowed Anthropic to restore limited Mythos access under tighter controls, while broader Fable access remains restricted.
- PreviewOpenAI previews GPT-5.6 LunaOpenAIsource ↗
Luna is the lower-cost GPT-5.6 preview variant, still gated by the government-review access process.
- PreviewOpenAI previews GPT-5.6 TerraOpenAIsource ↗
Terra is the mid-tier GPT-5.6 preview variant in the Sol/Terra/Luna rollout.
- UpdatedClaude Fable 5 remains restricted after Mythos carveoutAnthropicsource ↗
The June 26 government carveout restored only narrow Mythos access; Fable remains withdrawn from broad public availability.
- PreviewOpenAI begins limited GPT-5.6 preview under U.S. government reviewOpenAIsource ↗
The Sol, Terra, and Luna variants entered a tightly restricted preview for vetted customers while U.S. officials review security risks.
- UpdatedGemini 3.5 Pro launch reportedly slips toward JulyGoogle DeepMindsource ↗
Reporting says Google delayed broad Gemini 3.5 Pro release from June while testers continue using it in Antigravity and LMArena.
- ReleasedByteDance releases Seed 2.1 ProByteDance Seedsource ↗
ByteDance officially releases the Seed 2.1 family, led by Seed 2.1 Pro (Doubao-Seed-2.1-pro): a deep-thinking flagship agent model for the "coding and agent era" with a 256K context and strong image/video understanding. ByteDance positions its coding, agent, and multimodal capabilities as comparable to GPT-5.5, citing the highest score on GDPVal, top-tier Agents' Last Exam results, and SOTA visual/video-understanding benchmarks. Proprietary; served via Doubao and Volcano Engine.
- ReleasedByteDance releases Seed 2.1 TurboByteDance Seedsource ↗
ByteDance releases Seed 2.1 Turbo (Doubao-Seed-2.1-turbo) alongside Seed 2.1 Pro: a low-cost, low-latency tier built for large-scale production, feature-complete and positioned as performance-comparable to Pro, with a 256K context and list pricing roughly half that of the Pro tier. Proprietary; served via Doubao and Volcano Engine.
- Released
- ReleasedSakana AI releases FuguSakana AIsource ↗
- UpdatedZ.ai founder hints at a Fable-class frontier modelZ.ai (Zhipu AI)source ↗
Jie Tang said a Chinese Fable 5-class model would arrive sooner than Elon Musk's Q1 prediction; no product name or launch date has been confirmed.
- ReleasedMoonshot releases Kimi K2.7 CodeMoonshot AIsource ↗
Open coding-focused Kimi model with 1T total / 32B active parameters, native image/video input, and always-on thinking mode.
- ReleasedZ.ai releases GLM-5.2 with a 1M-token contextZ.ai (Zhipu AI)source ↗
MIT-licensed GLM flagship focused on long-horizon coding, agentic engineering, and IndexShare sparse-attention reuse.
- ReleasedMiniMax releases MiniMax-M3MiniMaxsource ↗
Native multimodal 428B/23B-active model with one-million-token context and MiniMax Sparse Attention.
- WithdrawnFable 5 access suspended after U.S. government interventionAnthropicsource ↗
Access to Fable and Mythos was suspended after government action tied to cybersecurity concerns.
- WithdrawnUS government orders Anthropic to pull Fable 5 and Mythos 5Anthropicsource ↗
Access suspended three days after launch under an export-control directive citing national security.
- ReleasedGoogle DeepMind releases DiffusionGemma 26B-A4BGoogle DeepMindsource ↗
First open-weight text-diffusion model at scale, built on the Gemma 4 26B-A4B MoE; denoises text in parallel 256-token blocks for up to ~4x faster generation. Apache-2.0.
- ReleasedCohere releases North Mini Code 1.0Coheresource ↗
Cohere's first developer-focused model and the first in its North code-agent family: a 30B/3B-active MoE for agentic coding with a 256K context. Apache-2.0.
- AnnouncedGPT-5.6 announced with a 1.5M-token context windowOpenAIsource ↗
OpenAI previewed GPT-5.6, claiming the largest context window of any frontier model.
- ReleasedClaude Fable 5 released as a Mythos-class modelAnthropicsource ↗
Anthropic launched Fable 5 as a safeguarded Mythos-class model for broader use.
- ReleasedClaude Fable 5 released — first public Mythos-class modelAnthropicsource ↗
Available across the Claude API, AWS, and Microsoft Foundry.
- ReleasedUnisound releases U2 native agentic large modelUnisoundsource ↗
Unisound launches U2, a general-purpose "native agentic" LLM built for execution that can autonomously decompose and complete 100+ step workflows; it reports 87.9 on GPQA Diamond and 75 on SWE-bench Verified with ~25% lower thinking-token use. Available on the Unisound Token Hub.
- ReleasedNVIDIA releases Nemotron 3 Ultra 550B-A55BNVIDIAsource ↗
Largest Nemotron 3 model appears on NVIDIA NIM with downloadable weights, 1M context, and agentic reasoning positioning.
- ReleasedGoogle DeepMind releases Gemma 4 12BGoogle DeepMindsource ↗
Dense 12B with a unified, encoder-free multimodal architecture and native audio input; runs on a 16GB laptop. Apache-2.0, 256K context.
- UpdatedOpenAI updates GPT-Rosalind with GPT-5.5 capabilitiesOpenAIsource ↗
A new GPT-Rosalind update folds in GPT-5.5's agentic coding and tool use, improving medicinal-chemistry and genomics performance while using ~31% fewer tokens; expanded to eligible organizations globally via trusted access.
- ReleasedMicrosoft unveils MAI-Thinking-1 at Build 2026Microsoftsource ↗
Microsoft's first in-house frontier reasoning model: a sparse MoE (~35B active, 256K context) trained without third-party distillation; 97.0% on AIME 2025.
- ReleasedNex AGI open-sources Nex-N2-ProNex AGIsource ↗
Shanghai Innovation Institute's Nex alliance releases Nex-N2-Pro, an open-weight (Apache-2.0) agentic model post-trained on Qwen3.5-397B-A17B (397B total / ~17B active) with an "Agentic Thinking" framework, ~262K context, and text+image input.
- ReleasedMicrosoft releases MAI-Code-1-FlashMicrosoftsource ↗
An efficient ~5B-active agentic coding model rolling out in GitHub Copilot; Microsoft reports a +16-point SWE-bench Pro lead over Claude Haiku 4.5.
- ReleasedAlibaba releases Qwen3.7-Plus multimodal agent modelAlibaba (Qwen)source ↗
Lower-cost multimodal sibling of Qwen3.7-Max with text, image, and video input and a 1M-token context.
May 2026
- ReleasedStepFun releases Step-3.7-FlashStepFunsource ↗
High-efficiency multimodal sparse-MoE (~196B/11B-active) vision-language model with a 256K context and selectable reasoning tiers, for coding agents and search workflows.
- ReleasedLiquid AI releases LFM2.5-8B-A1BLiquid AIsource ↗
Liquid AI ships its on-device MoE (8.3B total / ~1.5B active) with a 131K context that runs in under ~6GB of memory, under the LFM Open License, scaling pretraining to 38T tokens over the October 2025 LFM2-8B-A1B.
- ReleasedClaude Opus 4.8 releasedAnthropicsource ↗
Agentic upgrades and stronger long-running task performance.
- ReleasedMiniMax releases MiniMax-M2.7MiniMaxsource ↗
Open-weight agentic model focused on software engineering, productivity tasks, and model self-evolution workflows.
- ReleasedGoogle launches Gemini 3.5 Flash at I/O 2026Google DeepMindsource ↗
Fast, cost-efficient Gemini 3.5 tier with a 1M-token context and text/image/audio/video input; Google says it beats Gemini 3.1 Pro on coding and tool-use while running ~4x faster.
- AnnouncedGemini 3.5 Pro announced but delayed at Google I/O 2026Google DeepMindsource ↗
Google said its most powerful model in the works needed until the following month before release.
- ReleasedAlibaba launches Qwen3.7-Max, "The Agent Frontier"Alibaba (Qwen)source ↗
Proprietary agent-first flagship with a 1M-token context, OpenAI/Anthropic-compatible APIs, and long-horizon tool use.
- AnnouncedGemini 3.5 Pro announced at Google I/O 2026Google DeepMindsource ↗
- ReleasedQwen releases Qwen3.6-27BAlibaba (Qwen)source ↗
- RetiredClaude 3.7 Sonnet retiredAnthropicsource ↗
Endpoint shut down as part of Anthropic's 2026 deprecation calendar.
- ReleasedBaidu releases ERNIE 5.1Baidusource ↗
Sparse-MoE flagship; first Chinese model to reach the global LMArena top tier, reportedly trained at ~6% of comparable frontier pre-training cost.
- ReleasedOpenAI releases GPT-5.5OpenAIsource ↗
OpenAI releases GPT-5.5 with an 800K-token input context, 128K-token output limit, and stronger reasoning, coding, and multimodal performance.
- PreviewOpenAI previews GPT-5.5-Cyber for vetted defendersOpenAIsource ↗
Limited TAC preview for specialized authorized cybersecurity workflows, paired with stronger verification and safeguards.
- ReleasedGrok 4.3 generally availablexAIsource ↗
1M-token context at $1.25/$2.50 per million tokens; also live on Microsoft Foundry.
April 2026
- Released
- ReleasedDeepSeek V4-Pro preview released with 1M contextDeepSeeksource ↗
DeepSeek introduced the V4 preview series under MIT, led by V4-Pro at 1.6T total / 49B active parameters.
- ReleasedTencent open-sources Hunyuan Hy3-previewTencent Hunyuansource ↗
Tencent releases and open-sources the Hy3 preview, its rebuilt third-generation Hunyuan: a 295B-total / 21B-active MoE with a 256K context, positioned as a leading open reasoning-and-agent model for its size. Vendor-reported scores include 74.4 on SWE-bench Verified, 54.4 on Terminal-Bench 2.0, and 70.2 on WideSearch. Weights on GitHub and Hugging Face under Tencent's community license.
- ReleasedXiaomi open-sources MiMo-V2.5Xiaomi (MiMo)source ↗
Alongside the Pro flagship, Xiaomi releases MiMo-V2.5, a ~310B/15B-active sparse-MoE model trained on ~48T tokens with a 1M-token context, under the MIT license.
- ReleasedXiaomi open-sources MiMo-V2.5-ProXiaomi (MiMo)source ↗
Xiaomi releases its open-weight flagship MiMo-V2.5-Pro: a 1.02T-parameter MoE (~42B active) with hybrid attention and a 1M-token context, tuned for frontier-class agentic coding. MIT-licensed, with weights on Hugging Face.
- ReleasedTencent releases Hunyuan-A13B-InstructTencent Hunyuansource ↗
80B/13B-active Hunyuan MoE model released with open weights and agentic tool-use support.
- ReleasedOpenAI introduces GPT-Rosalind for life sciencesOpenAIsource ↗
OpenAI launches GPT-Rosalind, a frontier reasoning model purpose-built for biology, drug discovery, and translational medicine, available as a research preview in ChatGPT, Codex, and the API through a trusted-access program.
- ReleasedMeta releases Muse Spark for Meta AIMeta AIsource ↗
Meta releases Muse Spark, identified in reporting as the productized model behind Meta AI for U.S. users and the public successor to the Avocado effort.
- UpdatedMeta Avocado codename superseded by Muse Spark reportingMeta AIsource ↗
The Avocado rumor row is retained for provenance; the public model is tracked separately as Muse Spark.
- Released
- AnnouncedAnthropic discloses Claude Mythos but withholds public releaseAnthropicsource ↗
Frontier model shipped only to ~50 defensive-security partners via Project Glasswing.
- RetiredGPT-4o fully retired from ChatGPTOpenAIsource ↗
Removed from all ChatGPT plans after a Feb 13 deprecation notice.
- ReleasedGoogle DeepMind releases Gemma 4Google DeepMindsource ↗
Gemma 4 introduces advanced reasoning open models in 12B, 26B, and 31B sizes.
- ReleasedZ.ai launches GLM-5V-Turbo multimodal vision modelZ.ai (Zhipu AI)source ↗
Z.ai releases GLM-5V-Turbo, its first natively multimodal vision agent: image/video/text input with agent-oriented output (tool calling, task decomposition, GUI interaction) and a ~203K-token context.
March 2026
- ReleasedKimi K2.6 releasedMoonshot AIsource ↗
- ReleasedMistral Medium 3.5 releasedMistral AIsource ↗
- Released
- Released
- Released
- UpdatedReports say Meta delayed its Avocado modelMeta AIsource ↗
Meta's internal model reportedly missed a March target after underperforming top frontier systems; release timing and branding remain unconfirmed.
- ReleasedSarvam AI open-sources Sarvam-105BSarvam AIsource ↗
Apache-2.0 MoE model focused on reasoning, coding, agentic tasks, and Indian-language performance.
- Released
- Released
- ReleasedAlibaba releases the small Qwen3.5 family (0.8B-9B)Alibaba (Qwen)source ↗
Alibaba expands Qwen3.5 with four small dense models (9B, 4B, 2B, 0.8B) under Apache-2.0. The 9B is rated the most intelligent model under 10B parameters, with native vision and a 262K-token context.
- ReleasedQwen3.5-2B released for edge / on-device useAlibaba (Qwen)source ↗
- Released
February 2026
- Released
- ReleasedGemini 3.1 Pro generally availableGoogle DeepMindsource ↗
- ReleasedZ.ai releases GLM-5 for complex systems engineeringZ.ai (Zhipu AI)source ↗
- ReleasedOpenAI releases GPT-5.3-CodexOpenAIsource ↗
OpenAI updates Codex with GPT-5.3-Codex, improving code quality, repository-scale reasoning, and long-running agentic coding workflows.
- Released
- ReleasedQwen releases Qwen3-Coder-Next for coding agentsAlibaba (Qwen)source ↗
January 2026
- ReleasedMoonshot releases Kimi K2.5Moonshot AIsource ↗
Open multimodal K2 upgrade with MoonViT, thinking modes, visual coding, and agent-swarm workflows.
- Released
December 2025
- ReleasedOpenAI releases GPT-5.2-CodexOpenAIsource ↗
OpenAI releases GPT-5.2-Codex for software engineering agents, with a 400K-token input context and 128K-token output limit.
- Released
- ReleasedAi2 releases OLMo 3 Think 32B as a fully open reasoning modelAllen Institute for AI (Ai2)source ↗
- ReleasedOpenAI releases GPT-5.2OpenAIsource ↗
OpenAI releases GPT-5.2 as a stronger general GPT-5 model for reasoning, coding, vision, instruction following, and long-context analysis.
- Released
- Released
- Released
- Released
November 2025
- ReleasedLiquid AI releases LFM2 1.2BLiquid AIsource ↗
- ReleasedMoonshot releases Kimi K2 ThinkingMoonshot AIsource ↗
Open K2 reasoning-agent variant for deep thinking and stable long-horizon tool orchestration.
October 2025
- ReleasedMoonshot releases Kimi Linear 48B-A3BMoonshot AIsource ↗
MIT-licensed hybrid linear-attention checkpoints with a 1M-token context and lower KV-cache use.
- ReleasedAnthropic releases Claude Haiku 4.5Anthropicsource ↗
Anthropic releases Claude Haiku 4.5 as a faster, lower-cost Claude 4.5 tier for coding, tool-use, and latency-sensitive agents.
September 2025
- Released
- ReleasedAnthropic releases Claude Sonnet 4.5Anthropicsource ↗
Anthropic releases Claude Sonnet 4.5, positioning it as its strongest model for coding, agents, and computer-use workflows at launch.
- Released
- Updated
- ReleasedMoonshot updates Kimi K2 Instruct with 256K contextMoonshot AIsource ↗
The 0905 update improves agentic coding and frontend generation while doubling context length.
- ReleasedGoogle releases Gemma 3 (multimodal, 128k context)Google DeepMindsource ↗
August 2025
- Released
- Released
- UpdatedReports cite DeepSeek R2 delay tied to training hardwareDeepSeeksource ↗
R2 reportedly switched back to Nvidia for training after Huawei Ascend issues, with domestic hardware still targeted for inference.
- ReleasedZ.ai releases GLM-4.5V for multimodal reasoningZ.ai (Zhipu AI)source ↗
- UpdatedElon Musk says Grok 5 is planned before year-endxAIsource ↗
The statement is tracked as a rumor until a public Grok 5 model release is confirmed.
- ReleasedAnthropic releases Claude Opus 4.1Anthropicsource ↗
Anthropic releases Claude Opus 4.1 with improvements for coding, reasoning, and agentic reliability over Claude Opus 4.
- Released
- Released
- ReleasedGoogle releases Gemini 2.5 Deep ThinkGoogle DeepMindsource ↗
Google makes its more deliberative Gemini 2.5 Deep Think reasoning mode available after previewing it at Google I/O 2025.
July 2025
- ReleasedTII releases Falcon-H1 hybrid attention-SSM modelsTechnology Innovation Institutesource ↗
- Released
- ReleasedZ.ai releases GLM-4.5 under MITZ.ai (Zhipu AI)source ↗
- ReleasedQwen releases Qwen3-Coder-480B-A35B-InstructAlibaba (Qwen)source ↗
Alibaba Qwen releases its large open coding-agent MoE with 480B total / 35B active parameters and a 256K-token native context.
- ReleasedGoogle releases Gemini 2.5 Flash-LiteGoogle DeepMindsource ↗
Google releases Gemini 2.5 Flash-Lite as the lowest-cost, lowest-latency Gemini 2.5 tier for high-volume production tasks.
- ReleasedLG AI Research releases EXAONE 4.0 32BLG AI Researchsource ↗
- ReleasedMoonshot releases Kimi K2 InstructMoonshot AIsource ↗
Original open 1T-parameter K2 MoE release optimized for coding, reasoning, and agentic tool use.
- Released
- ReleasedHugging Face releases SmolLM3 3BHugging Facesource ↗
June 2025
- Released
- Released
- ReleasedMoonshot releases Kimi-VL-A3B-Thinking-2506Moonshot AIsource ↗
Updated efficient multimodal reasoning model with stronger video, high-resolution perception, and lower thinking-token use.
- ReleasedGoogle releases Gemini 2.5 FlashGoogle DeepMindsource ↗
Google brings Gemini 2.5 reasoning improvements to a faster, lower-cost Flash production tier with a 1M-token context.
- ReleasedMoonshot releases Kimi-Dev-72BMoonshot AIsource ↗
Open coding LLM trained with repository-level reinforcement learning for issue resolution.
- Released
- ReleasedMistral releases Magistral SmallMistral AIsource ↗
- Released
May 2025
- Released
- ReleasedByteDance Seed releases Seed Thinking v1.5ByteDance Seedsource ↗
- Released
- PreviewMistral previews Devstral Small 2505Mistral AIsource ↗
Mistral and All Hands AI release Devstral Small 2505, a 24B Apache-2.0 coding-agent model for repository-level software engineering tasks.
- ReleasedSarvam AI releases Sarvam-MSarvam AIsource ↗
- ReleasedGoogle releases Gemma 3n E4BGoogle DeepMindsource ↗
Google releases the mobile-first Gemma 3n E4B variant for efficient on-device multimodal inference.
- ReleasedMistral releases Mistral Medium 3Mistral AIsource ↗
Mistral releases Medium 3 as a lower-cost enterprise workhorse for coding, STEM, search, and multilingual workloads.
April 2025
- Released
- Released
- ReleasedQwen releases Qwen3-235B-A22BAlibaba (Qwen)source ↗
- ReleasedMoonshot releases Kimi-Audio-7B-InstructMoonshot AIsource ↗
Open audio foundation model for speech recognition, audio QA, captioning, generation, and conversation.
- ReleasedMoonshot releases Kimi-VL-A3B-InstructMoonshot AIsource ↗
Efficient MIT-licensed vision-language MoE for OCR, video, long documents, and agent tasks.
- Released
- Released
- AnnouncedMeta previews Llama 4 Behemoth but does not release weightsMeta AIsource ↗
Meta described Behemoth as a still-training 288B-active / nearly 2T-total teacher model used to distill Llama 4 Scout and Maverick.
- Released
- Released
- Released
March 2025
- ReleasedQwen releases Qwen2.5-Omni-7BAlibaba (Qwen)source ↗
- ReleasedGoogle DeepMind releases Gemini 2.5 ProGoogle DeepMindsource ↗
- Released
- Released
- Released
- Released
- ReleasedAi2 releases OLMo 2 32B — fully open weights, data, and codeAllen Institute for AI (Ai2)source ↗
February 2025
- Released
- ReleasedMoonshot releases Moonlight-16B-A3B-InstructMoonshot AIsource ↗
Open 16B/3B-active MoE demonstrating Moonshot's scalable Muon optimizer work.
- Released
- Released
- ReleasedCognitive Computations releases Dolphin 3.0 Llama 3.1 8BCognitive Computationssource ↗
January 2025
- Released
- ReleasedQwen releases Qwen2.5-MaxAlibaba (Qwen)source ↗
- ReleasedQwen releases Qwen2.5-VL-72BAlibaba (Qwen)source ↗
- ReleasedByteDance Seed releases Doubao-1.5-proByteDance Seedsource ↗
- ReleasedMoonshot releases Kimi k1.5Moonshot AIsource ↗
Multimodal reinforcement-learning reasoning model reported to match OpenAI o1 on math, coding, and multimodal reasoning.
- Released
- Released
December 2024
- Released
- Released
- Released
- ReleasedTII releases the Falcon 3 small-model familyTechnology Innovation Institutesource ↗
- Released
- Released
- ReleasedGoogle DeepMind releases Gemini 2.0 FlashGoogle DeepMindsource ↗
- ReleasedLG AI Research releases EXAONE 3.5 32BLG AI Researchsource ↗
- Released
- Released
- Released
- Released
November 2024
- ReleasedQwen releases QwQ-32B-PreviewAlibaba (Qwen)source ↗
- ReleasedAi2 releases Tulu 3 405BAllen Institute for AI (Ai2)source ↗
- Released
- ReleasedQwen releases Qwen2.5-Coder-32BAlibaba (Qwen)source ↗
- ReleasedHugging Face releases SmolLM2 1.7BHugging Facesource ↗
- Released
October 2024
- ReleasedSarvam AI releases Sarvam-1Sarvam AIsource ↗
- Released
- Released
- ReleasedMistral releases Ministral 8BMistral AIsource ↗
- Released
- Released
September 2024
- Released
- ReleasedAi2 releases Molmo 72BAllen Institute for AI (Ai2)source ↗
- ReleasedQwen releases Qwen2.5-72BAlibaba (Qwen)source ↗
- ReleasedMistral releases Pixtral 12BMistral AIsource ↗
- Released
- Released
- ReleasedTencent Hunyuan releases Hunyuan TurboTencent Hunyuansource ↗
- Released
- ReleasedAi2 releases OLMoE 1B-7BAllen Institute for AI (Ai2)source ↗
August 2024
- Released
- Released
- Released
- Released
- ReleasedLG AI Research releases EXAONE 3.0 7.8BLG AI Researchsource ↗
- Released
July 2024
- Released
- ReleasedMistral releases Mistral NeMoMistral AIsource ↗
June 2024
- ReleasedGoogle DeepMind releases Gemma 2 27BGoogle DeepMindsource ↗
- Released
- Released
- Released
- ReleasedQwen releases Qwen2-72BAlibaba (Qwen)source ↗
- ReleasedZhipu AI releases GLM-4-9BZ.ai (Zhipu AI)source ↗
May 2024
- ReleasedMistral releases Codestral 22BMistral AIsource ↗
- Released
- ReleasedByteDance Seed releases Doubao-proByteDance Seedsource ↗
- Released
- ReleasedTII releases Falcon 2 11BTechnology Innovation Institutesource ↗
- Released
- Released
April 2024
- Released
- ReleasedSnowflake releases Snowflake ArcticSnowflake AI Researchsource ↗
- Released
- Released
- Released
- ReleasedMistral releases Mixtral 8x22BMistral AIsource ↗
- Released
- Released
- ReleasedGoogle DeepMind releases CodeGemma 7BGoogle DeepMindsource ↗
- Released
March 2024
- ReleasedAI21 releases JambaAI21 Labssource ↗
- Released
- ReleasedDatabricks releases DBRX InstructDatabricks / MosaicMLsource ↗
- Released
- ReleasedMoonshot releases Kimi 1MMoonshot AIsource ↗
- Released
- Released
February 2024
- Released
- ReleasedMistral releases Mistral LargeMistral AIsource ↗
- ReleasedGoogle DeepMind releases Gemma 7BGoogle DeepMindsource ↗
- ReleasedGoogle DeepMind releases Gemini 1.5 ProGoogle DeepMindsource ↗
- ReleasedQwen releases Qwen1.5-110BAlibaba (Qwen)source ↗
- ReleasedQwen releases Qwen1.5-72B-ChatAlibaba (Qwen)source ↗
Qwen releases the 72B chat-tuned Qwen1.5 checkpoint with 32K context and improved alignment.
- ReleasedAi2 releases OLMo 7BAllen Institute for AI (Ai2)source ↗
January 2024
- ReleasedStability AI releases Stable LM 2 1.6BStability AIsource ↗
- ReleasedZhipu AI releases GLM-4Z.ai (Zhipu AI)source ↗
- ReleasedNous Research releases Nous Hermes 2 MixtralNous Researchsource ↗
- Released
- Released
- Released
December 2023
- Released
- ReleasedMicrosoft releases Phi-2Microsoftsource ↗
- Released
- ReleasedGoogle announces Gemini 1.0Google DeepMindsource ↗
November 2023
- ReleasedAlibaba open-sources Qwen-72BAlibaba (Qwen)source ↗
- Released
- Released01.AI releases Yi-34B-Chat01.AIsource ↗
01.AI releases the chat-tuned Yi-34B checkpoint alongside quantized chat variants.
- Released
- Released
- Released
- Released
- Released
October 2023
- Released
- ReleasedMoonshot releases Kimi ChatMoonshot AIsource ↗
September 2023
- Released
- Released
- ReleasedMistral AI releases Mistral 7BMistral AIsource ↗
- ReleasedQwen releases Qwen-14BAlibaba (Qwen)source ↗
- ReleasedTencent Hunyuan releases HunyuanTencent Hunyuansource ↗
- Released
- ReleasedTII releases Falcon 180BTechnology Innovation Institutesource ↗
August 2023
- Released
- ReleasedQwen releases Qwen-7BAlibaba (Qwen)source ↗
July 2023
- Released
- ReleasedLG AI Research releases EXAONE 2.0LG AI Researchsource ↗
- Released
- Released
June 2023
- ReleasedZhipu AI releases ChatGLM2-6BZ.ai (Zhipu AI)source ↗
- ReleasedMicrosoft releases Phi-1Microsoftsource ↗
May 2023
- ReleasedTII releases Falcon 40BTechnology Innovation Institutesource ↗
- ReleasedGoogle DeepMind releases PaLM 2Google DeepMindsource ↗
- ReleasedDatabricks releases MPT-7BDatabricks / MosaicMLsource ↗
March 2023
- ReleasedLMSYS releases Vicuna 13BLMSYS / SkyLabsource ↗
- Released
- ReleasedAnthropic releases Claude 1Anthropicsource ↗
- ReleasedZhipu AI releases ChatGLM-6BZ.ai (Zhipu AI)source ↗
- Released
- Released
- Released
February 2023
- Released
November 2022
- UpdatedChatGPT (GPT-3.5) launches and reaches 100M usersOpenAIsource ↗
The consumer launch that ignited the modern LLM race.
- WithdrawnMeta pulls the Galactica demo after three daysMeta AIsource ↗
Withdrawn following criticism of confident but inaccurate scientific output.
- Released
July 2022
- Released
April 2022
- ReleasedGoogle announces PaLM (540B)Google DeepMindsource ↗
December 2021
- ReleasedLG AI Research releases EXAONE 1.0LG AI Researchsource ↗
- Released
August 2021
- Released
June 2020
- Released
November 2019
- ReleasedOpenAI releases the full GPT-2 (1.5B) weightsOpenAIsource ↗
After a staged rollout that began in Feb 2019 over misuse concerns.
October 2018
- ReleasedGoogle releases BERTGoogle DeepMindsource ↗