LLM Releases

Changelog

Last updated Jul 26, 2026

Everything, in order

A single feed of releases, updates, deprecations, and retractions across every lab we track. Each item links to the model and its primary source.

July 2026

  1. Updated
    Moonshot publishes Kimi K3 open weights
    Moonshot AIsource ↗

    Moonshot AI publishes the full Kimi K3 weights to Hugging Face under a Modified MIT license on July 26 — a day ahead of its announced July 27 target — making the 2.8T-parameter MoE freely downloadable, modifiable, and self-hostable, and cementing K3 as the largest open-weight model publicly available.

  2. Released
    Anthropic releases Claude Opus 5
    Anthropicsource ↗

    Anthropic releases Claude Opus 5, its flagship model for demanding reasoning, autonomous coding, and long-horizon agentic work — pitched as the go-to model for most knowledge work, approaching Fable 5 capability in many categories at about half the price. Adds a five-level 'effort' dial on the Claude API/Platform to trade compute for capability, a 1M-token context at standard pricing, and up to 128K output tokens. Standard pricing $5/$25 per Mtok (matching Opus 4.8), plus a $10/$50 Fast mode; becomes the default for Claude Max subscribers. Anthropic calls it its most aligned Opus model.

  3. Released
    Ant Group's inclusionAI releases Ling-3.0-flash
    Ant Group (inclusionAI)source ↗

    Ant Group's inclusionAI lab releases Ling-3.0-flash, a hybrid-reasoning Mixture-of-Experts model with 124B total parameters and ~5.1B active per token (1/64 expert activation), built for production-scale agents. Ant claims it matches or beats its own ~1T-parameter flagship on most benchmarks shown at 1/8 the total and 1/12 the active parameters — a vendor claim with no public benchmark table at launch. Uses a KDA + MLA hybrid-linear attention stack at a reported 5:1 ratio for an economical 256K-token context. Announced as open-weight under Apache 2.0, but weights and a model card were not yet posted to Hugging Face as of July 24; usable only via hosted API, free on OpenRouter and Vercel AI Gateway through August 3 2026.

  4. Released
    Google releases Gemini 3.5 Flash-Lite
    Google DeepMindsource ↗

    Fastest, most cost-effective 3.5-class model (~350 output tokens/s) for high-throughput agentic workloads, priced at $0.30 / $2.50 per Mtok with a 1M-token context and built-in computer use. Large step up on 3.1 Flash-Lite and beats 3 Flash on several agentic/coding evals.

  5. Released
    Poolside releases Laguna S 2.1
    Poolsidesource ↗

    Poolside releases Laguna S 2.1, a 118B-total / 8B-active open-weight MoE coding model with a 1M-token context, pitched as 'the West's most capable open-weight model' for its weight class. It scores 70.2% on Terminal-Bench 2.1 and tops the published open disclosed-size table on SWE-bench Multilingual at 78.5%. Trained in under nine weeks on 4,096 H200 GPUs, it ships weights on Hugging Face under OpenMDW-1.1 (BF16/FP8/INT4/NVFP4 + GGUF/MLX) and runs at 4-bit on a single NVIDIA DGX Spark. Hosted free at 256K context and paid at full 1M context via OpenRouter ($0.10/$0.20/$0.01 per 1M input/output/cache-read tokens).

  6. Preview
    Google previews Gemini 3.5 Flash Cyber in CodeMender
    Google DeepMindsource ↗

    A cyber-specialized model built on 3.5 Flash for finding and fixing software vulnerabilities, deployed inside the CodeMender agent and reaching competitive frontier performance on CyberGym. Limited-access pilot restricted to governments and trusted partners.

  7. Released
    Google launches Gemini 3.6 Flash
    Google DeepMindsource ↗

    Google DeepMind ships Gemini 3.6 Flash, a multimodal 1M-context workhorse for agentic workflows that improves on 3.5 Flash in coding, knowledge work, and computer use while cutting output-token usage ~17% at a lower price ($1.50 / $7.50 per Mtok). Available in the Gemini API, Gemini Enterprise, and the Gemini app.

  8. Updated
    DeepSeek V4-Flash reaches general availability
    DeepSeeksource ↗

    V4-Flash, the efficient tier of the V4 family, exits preview alongside V4-Pro as DeepSeek retires its legacy deepseek-chat and deepseek-reasoner aliases (cutoff July 24, 2026).

  9. Updated
    DeepSeek V4 family reaches general availability
    DeepSeeksource ↗

    DeepSeek moves the V4 model family (V4-Pro and V4-Flash) out of preview and into general availability, closing a run of just under three months from the April 24 preview. The legacy deepseek-chat and deepseek-reasoner API aliases are retired on July 24, 2026 with no fallback; production traffic must reference deepseek-v4-pro / deepseek-v4-flash.

  10. Preview
    Alibaba unveils Qwen3.8-Max-Preview
    Alibaba (Qwen)source ↗

    Alibaba launches Qwen3.8, a 2.4-trillion-parameter fully-multimodal model — more than double the size of its predecessor — which the Qwen team says ranks "second only to Fable 5" on overall performance (internal evals, no independent benchmarks yet). Qwen3.8-Max-Preview is available to developers via Alibaba's Token Plan subscription and the Qoder / QoderWork coding platforms, with open weights promised "soon" but no timeline, license, or architecture details disclosed.

  11. Released
    Moonshot AI releases Kimi K3
    Moonshot AIsource ↗

    Largest open-weight model to date: 2.8T-parameter MoE (896 experts, 16 active) with Kimi Delta Attention, native multimodal input, and a 1M-token context. API launched Jul 16 at $3/$15 per Mtok; full weights slated for Jul 27. Reported 93.5% GPQA Diamond (strongest published open-weight result).

  12. Updated
    Gemini 3.5 Pro slips past its July 17 target
    Google DeepMindsource ↗

    Bloomberg reports Google delayed Gemini 3.5 Pro again after the rebuilt model fell short of internal quality goals on hallucinations and reliability — its third slipped target after June and early July. Google DeepMind's Logan Kilpatrick said on July 21 the company is testing it with partners and hopes to "land soon"; there is still no model card or public API entry.

  13. Released
    Thinking Machines releases Inkling
    Thinking Machines Labsource ↗

    Mira Murati's lab ships its first model — a 975B-total / 41B-active multimodal MoE (text/image/audio in, text out) pretrained on ~45T tokens, released under Apache-2.0 with weights on Hugging Face (1M context) and hosted on the Tinker API (256K). Debuts at 41 on the Artificial Analysis Intelligence Index, the leading U.S. open-weights model.

  14. Released
    Kwaipilot ships KAT-Coder-Air V2.5
    Kwaipilot (Kuaishou)source ↗

    The efficient ~32B-active variant of KAT-Coder V2.5, sharing the 256K context and agentic tool-use focus at roughly a fifth of Pro's price ($0.15/$0.60 per Mtok).

  15. Released
    Kwaipilot ships KAT-Coder-Pro V2.5
    Kwaipilot (Kuaishou)source ↗

    Kuaishou's Kwaipilot team releases KAT-Coder-Pro V2.5, an agentic coding MoE (~72B active) trained with large-scale agentic RL in verifiable repository environments, with a 256K context, 80K max output, and API pricing of $0.74/$2.96 per Mtok via StreamLake, Atlas Cloud, and OpenRouter.

  16. Released
    GPT-5.6 Luna reaches general availability
    OpenAIsource ↗

    Luna, the fast, most cost-efficient tier of the GPT-5.6 family ($1/$6 per Mtok), moves from restricted preview to general availability alongside Sol and Terra.

  17. Released
    Meta launches Muse Spark 1.1 — its first paid model
    Meta AIsource ↗

    Meta Superintelligence Labs releases Muse Spark 1.1 in US public preview on the Meta Model API, the first time Meta charges for one of its models. A natively multimodal reasoning model (text/image/video/PDF/audio in, text out) with a self-compacting 1M-token context, aimed at agentic and coding workflows and priced at $1.25/$4.25 per Mtok — about a quarter of comparable Anthropic/OpenAI models. Closed weights; benchmarks around the Opus 4.8 / GPT-5.5 tier.

  18. Released
    OpenAI takes the GPT-5.6 family to GA
    OpenAIsource ↗

    Following U.S. government review and a two-week restricted preview, OpenAI rolls out the full GPT-5.6 family — Sol, Terra, and Luna — to general availability across ChatGPT (Plus, Pro, Business, Enterprise) and the API.

  19. Released
    GPT-5.6 Terra reaches general availability
    OpenAIsource ↗

    Terra, the balanced everyday-work tier of the GPT-5.6 family ($2.50/$15 per Mtok), moves from restricted preview to general availability alongside Sol and Luna, available to Plus, Pro, Business, and Enterprise users in ChatGPT and via the API.

  20. Released
    GPT-5.6 Sol reaches general availability
    OpenAIsource ↗

    OpenAI moves the GPT-5.6 family to general availability after the June 26 restricted preview. Sol, the flagship for difficult professional, coding, research, computer-use, and tool-heavy work, ships with a 1.05M-token context window, 128K max output, and standard pricing of $5/$30 per Mtok, alongside a new max reasoning effort and ultra subagent mode.

  21. Released
    SpaceXAI and Cursor launch Grok 4.5
    xAIsource ↗

    SpaceXAI (the rebranded xAI) releases its most capable model to date — the first Grok trained jointly with Cursor (Anysphere), pitched as "Opus-class" but faster, cheaper ($2/$6 per Mtok), and ~4x more token-efficient than Opus 4.8 on SWE-bench Pro. Available in Grok Build (default), Cursor (all plans), and the SpaceXAI console; initially unavailable in the EU.

  22. Announced
    Mistral confirms a frontier-gap open-weight MoE with July early access
    Mistral AIsource ↗

    CEO Arthur Mensch confirms Mistral is preparing a new open-weight Mixture-of-Experts family — "fat but sparse" — entering early access in July 2026 and aimed at the frontier open-weight tier. No parameter count, benchmarks, license terms, or release date disclosed; tracked as rumored.

  23. Released
    NVIDIA releases Nemotron-Labs-3-Puzzle-75B-A9B
    NVIDIAsource ↗

    A compressed variant of Nemotron-3-Super produced with "Iterative Puzzle", trimming the parent to 75.3B total / 9.3B active while keeping the hybrid Mamba-Transformer LatentMoE design and 1M context. NVIDIA reports ~2x higher server throughput at matched user throughput; released on Hugging Face in BF16/FP8/NVFP4 under OpenMDW-1.1.

  24. Released
    Tencent officially releases and open-sources Hunyuan Hy3
    Tencent Hunyuansource ↗

    Tencent officially launches Hunyuan 3.0 (Hy3), the GA of its rebuilt third-generation model, and open-sources it under Apache-2.0. A 295B-total / 21B-active MoE (plus a 3.8B multi-token-prediction layer) with a 256K context and three selectable fast/slow inference modes; Tencent reports it rivals GLM-5.2 and DeepSeek-V4 and matches or surpasses GPT-5.5 on several science benchmarks, with 78.0 on SWE-bench Verified. Weights on Hugging Face (tencent/Hy3) and ModelScope, with a free OpenRouter route (tencent/hy3:free) through July 21, 2026.

  25. Released
    Poolside releases Laguna XS 2.1
    Poolsidesource ↗

    Poolside releases Laguna XS 2.1, an upgraded 33B-A3B open-weight MoE coding model served at 256K context. It raises SWE-bench Multilingual by 5.4 points to 63.1% over XS.2, adds open-weighted DFlash speculator models that roughly double local tokens/sec, and moves to the fully permissive OpenMDW-1.1 license. Weights are on Hugging Face (BF16/FP8/INT4/NVFP4) and it's available free on OpenRouter, with paid API pricing of $0.10/$0.20/$0.05 per 1M input/output/cache-read tokens.

  26. Updated
    Claude Fable 5 access restored globally
    Anthropicsource ↗

    Anthropic says Claude Fable 5 is available globally again on Claude Platform, Claude.ai, Claude Code, and Claude Cowork after export controls were lifted; cloud partner access is being re-enabled and safeguards now route high-risk requests to Opus 4.8.

  27. Updated
    Claude Mythos 5 access restored for approved partners
    Anthropicsource ↗

    Anthropic restored Mythos 5 access for approved U.S. organizations and continues expanding the Glasswing trusted-access program, while Mythos remains restricted rather than generally available.

June 2026

  1. Released
    Meituan open-sources LongCat-2.0, trained entirely on Chinese chips
    Meituan (LongCat)source ↗

    Meituan releases LongCat-2.0, a 1.6T-parameter MoE (~48B active) with a 1M-token context for agentic coding, open-sourced under the MIT license on Hugging Face and GitHub. It was trained and served on a ~50,000-card cluster of domestic Chinese AI chips — the first trillion-parameter model Meituan says completed full-process training and inference on home-grown hardware. Vendor-reported results: 59.5 SWE-bench Pro (vs GPT-5.5's 58.6), 70.8 Terminal-Bench 2.1, and 77.3 SWE-bench Multilingual.

  2. Released
    Anthropic releases Claude Sonnet 5
    Anthropicsource ↗

    Anthropic releases Claude Sonnet 5, its most agentic Sonnet yet — performance approaching Opus 4.8 at a lower price, made the default model on the Free and Pro plans and available to Max, Team, and Enterprise users. Available in Claude Code and via the Claude API as claude-sonnet-5, with introductory pricing of $2 per Mtok input / $10 per Mtok output through Aug 31, 2026, then standard $3/$15.

  3. Updated
    White House lifts Anthropic Fable 5 ban
    Anthropicsource ↗

    Reporting says the White House lifted the ban on Anthropic's models after an agreement on additional safeguards, authorizing Anthropic to return Fable 5 to public release channels while Mythos remains limited to pre-vetted partners.

  4. Released
    Base44 rolls out Base1, its first in-house model
    Base44source ↗

    Base44 (a Wix company) rolls out Base1, a general-purpose 'vibe coding' agent fine-tuned on an open-source foundation model using data from tens of millions of platform interactions, selectable alongside GPT-5.5 and Claude Opus 4.8 in its model picker.

  5. Preview
    OpenAI previews GPT-5.6 Sol
    OpenAIsource ↗

    Sol is the highest-capability GPT-5.6 preview tier, available only to a small vetted cohort.

  6. Updated
    Commerce Department clears limited Claude Mythos return
    Anthropicsource ↗

    U.S. Commerce officials reportedly allowed Anthropic to restore limited Mythos access under tighter controls, while broader Fable access remains restricted.

  7. Preview
    OpenAI previews GPT-5.6 Luna
    OpenAIsource ↗

    Luna is the lower-cost GPT-5.6 preview variant, still gated by the government-review access process.

  8. Preview
    OpenAI previews GPT-5.6 Terra
    OpenAIsource ↗

    Terra is the mid-tier GPT-5.6 preview variant in the Sol/Terra/Luna rollout.

  9. Updated
    Claude Fable 5 remains restricted after Mythos carveout
    Anthropicsource ↗

    The June 26 government carveout restored only narrow Mythos access; Fable remains withdrawn from broad public availability.

  10. Preview
    OpenAI begins limited GPT-5.6 preview under U.S. government review
    OpenAIsource ↗

    The Sol, Terra, and Luna variants entered a tightly restricted preview for vetted customers while U.S. officials review security risks.

  11. Updated
    Gemini 3.5 Pro launch reportedly slips toward July
    Google DeepMindsource ↗

    Reporting says Google delayed broad Gemini 3.5 Pro release from June while testers continue using it in Antigravity and LMArena.

  12. Released
    ByteDance releases Seed 2.1 Pro
    ByteDance Seedsource ↗

    ByteDance officially releases the Seed 2.1 family, led by Seed 2.1 Pro (Doubao-Seed-2.1-pro): a deep-thinking flagship agent model for the "coding and agent era" with a 256K context and strong image/video understanding. ByteDance positions its coding, agent, and multimodal capabilities as comparable to GPT-5.5, citing the highest score on GDPVal, top-tier Agents' Last Exam results, and SOTA visual/video-understanding benchmarks. Proprietary; served via Doubao and Volcano Engine.

  13. Released
    ByteDance releases Seed 2.1 Turbo
    ByteDance Seedsource ↗

    ByteDance releases Seed 2.1 Turbo (Doubao-Seed-2.1-turbo) alongside Seed 2.1 Pro: a low-cost, low-latency tier built for large-scale production, feature-complete and positioned as performance-comparable to Pro, with a 256K context and list pricing roughly half that of the Pro tier. Proprietary; served via Doubao and Volcano Engine.

  14. Released
  15. Released
  16. Updated
    Z.ai founder hints at a Fable-class frontier model
    Z.ai (Zhipu AI)source ↗

    Jie Tang said a Chinese Fable 5-class model would arrive sooner than Elon Musk's Q1 prediction; no product name or launch date has been confirmed.

  17. Released
    Moonshot releases Kimi K2.7 Code
    Moonshot AIsource ↗

    Open coding-focused Kimi model with 1T total / 32B active parameters, native image/video input, and always-on thinking mode.

  18. Released
    Z.ai releases GLM-5.2 with a 1M-token context
    Z.ai (Zhipu AI)source ↗

    MIT-licensed GLM flagship focused on long-horizon coding, agentic engineering, and IndexShare sparse-attention reuse.

  19. Released
    MiniMax releases MiniMax-M3
    MiniMaxsource ↗

    Native multimodal 428B/23B-active model with one-million-token context and MiniMax Sparse Attention.

  20. Withdrawn
    Fable 5 access suspended after U.S. government intervention
    Anthropicsource ↗

    Access to Fable and Mythos was suspended after government action tied to cybersecurity concerns.

  21. Withdrawn
    US government orders Anthropic to pull Fable 5 and Mythos 5
    Anthropicsource ↗

    Access suspended three days after launch under an export-control directive citing national security.

  22. Released
    Google DeepMind releases DiffusionGemma 26B-A4B
    Google DeepMindsource ↗

    First open-weight text-diffusion model at scale, built on the Gemma 4 26B-A4B MoE; denoises text in parallel 256-token blocks for up to ~4x faster generation. Apache-2.0.

  23. Released
    Cohere releases North Mini Code 1.0
    Coheresource ↗

    Cohere's first developer-focused model and the first in its North code-agent family: a 30B/3B-active MoE for agentic coding with a 256K context. Apache-2.0.

  24. Announced
    GPT-5.6 announced with a 1.5M-token context window
    OpenAIsource ↗

    OpenAI previewed GPT-5.6, claiming the largest context window of any frontier model.

  25. Released
    Claude Fable 5 released as a Mythos-class model
    Anthropicsource ↗

    Anthropic launched Fable 5 as a safeguarded Mythos-class model for broader use.

  26. Released
    Claude Fable 5 released — first public Mythos-class model
    Anthropicsource ↗

    Available across the Claude API, AWS, and Microsoft Foundry.

  27. Released
    Unisound releases U2 native agentic large model
    Unisoundsource ↗

    Unisound launches U2, a general-purpose "native agentic" LLM built for execution that can autonomously decompose and complete 100+ step workflows; it reports 87.9 on GPQA Diamond and 75 on SWE-bench Verified with ~25% lower thinking-token use. Available on the Unisound Token Hub.

  28. Released
    NVIDIA releases Nemotron 3 Ultra 550B-A55B
    NVIDIAsource ↗

    Largest Nemotron 3 model appears on NVIDIA NIM with downloadable weights, 1M context, and agentic reasoning positioning.

  29. Released
    Google DeepMind releases Gemma 4 12B
    Google DeepMindsource ↗

    Dense 12B with a unified, encoder-free multimodal architecture and native audio input; runs on a 16GB laptop. Apache-2.0, 256K context.

  30. Updated
    OpenAI updates GPT-Rosalind with GPT-5.5 capabilities
    OpenAIsource ↗

    A new GPT-Rosalind update folds in GPT-5.5's agentic coding and tool use, improving medicinal-chemistry and genomics performance while using ~31% fewer tokens; expanded to eligible organizations globally via trusted access.

  31. Released
    Microsoft unveils MAI-Thinking-1 at Build 2026
    Microsoftsource ↗

    Microsoft's first in-house frontier reasoning model: a sparse MoE (~35B active, 256K context) trained without third-party distillation; 97.0% on AIME 2025.

  32. Released
    Nex AGI open-sources Nex-N2-Pro
    Nex AGIsource ↗

    Shanghai Innovation Institute's Nex alliance releases Nex-N2-Pro, an open-weight (Apache-2.0) agentic model post-trained on Qwen3.5-397B-A17B (397B total / ~17B active) with an "Agentic Thinking" framework, ~262K context, and text+image input.

  33. Released
    Microsoft releases MAI-Code-1-Flash
    Microsoftsource ↗

    An efficient ~5B-active agentic coding model rolling out in GitHub Copilot; Microsoft reports a +16-point SWE-bench Pro lead over Claude Haiku 4.5.

  34. Released
    Alibaba releases Qwen3.7-Plus multimodal agent model
    Alibaba (Qwen)source ↗

    Lower-cost multimodal sibling of Qwen3.7-Max with text, image, and video input and a 1M-token context.

May 2026

  1. Released
    StepFun releases Step-3.7-Flash
    StepFunsource ↗

    High-efficiency multimodal sparse-MoE (~196B/11B-active) vision-language model with a 256K context and selectable reasoning tiers, for coding agents and search workflows.

  2. Released
    Liquid AI releases LFM2.5-8B-A1B
    Liquid AIsource ↗

    Liquid AI ships its on-device MoE (8.3B total / ~1.5B active) with a 131K context that runs in under ~6GB of memory, under the LFM Open License, scaling pretraining to 38T tokens over the October 2025 LFM2-8B-A1B.

  3. Released
    Claude Opus 4.8 released
    Anthropicsource ↗

    Agentic upgrades and stronger long-running task performance.

  4. Released
    MiniMax releases MiniMax-M2.7
    MiniMaxsource ↗

    Open-weight agentic model focused on software engineering, productivity tasks, and model self-evolution workflows.

  5. Released
    Google launches Gemini 3.5 Flash at I/O 2026
    Google DeepMindsource ↗

    Fast, cost-efficient Gemini 3.5 tier with a 1M-token context and text/image/audio/video input; Google says it beats Gemini 3.1 Pro on coding and tool-use while running ~4x faster.

  6. Announced
    Gemini 3.5 Pro announced but delayed at Google I/O 2026
    Google DeepMindsource ↗

    Google said its most powerful model in the works needed until the following month before release.

  7. Released
    Alibaba launches Qwen3.7-Max, "The Agent Frontier"
    Alibaba (Qwen)source ↗

    Proprietary agent-first flagship with a 1M-token context, OpenAI/Anthropic-compatible APIs, and long-horizon tool use.

  8. Announced
  9. Released
    Qwen releases Qwen3.6-27B
    Alibaba (Qwen)source ↗
  10. Retired
    Claude 3.7 Sonnet retired
    Anthropicsource ↗

    Endpoint shut down as part of Anthropic's 2026 deprecation calendar.

  11. Released
    Baidu releases ERNIE 5.1
    Baidusource ↗

    Sparse-MoE flagship; first Chinese model to reach the global LMArena top tier, reportedly trained at ~6% of comparable frontier pre-training cost.

  12. Released
    OpenAI releases GPT-5.5
    OpenAIsource ↗

    OpenAI releases GPT-5.5 with an 800K-token input context, 128K-token output limit, and stronger reasoning, coding, and multimodal performance.

  13. Preview
    OpenAI previews GPT-5.5-Cyber for vetted defenders
    OpenAIsource ↗

    Limited TAC preview for specialized authorized cybersecurity workflows, paired with stronger verification and safeguards.

  14. Released
    Grok 4.3 generally available
    xAIsource ↗

    1M-token context at $1.25/$2.50 per million tokens; also live on Microsoft Foundry.

April 2026

  1. Released
  2. Released
    DeepSeek V4-Pro preview released with 1M context
    DeepSeeksource ↗

    DeepSeek introduced the V4 preview series under MIT, led by V4-Pro at 1.6T total / 49B active parameters.

  3. Released
    Tencent open-sources Hunyuan Hy3-preview
    Tencent Hunyuansource ↗

    Tencent releases and open-sources the Hy3 preview, its rebuilt third-generation Hunyuan: a 295B-total / 21B-active MoE with a 256K context, positioned as a leading open reasoning-and-agent model for its size. Vendor-reported scores include 74.4 on SWE-bench Verified, 54.4 on Terminal-Bench 2.0, and 70.2 on WideSearch. Weights on GitHub and Hugging Face under Tencent's community license.

  4. Released
    Xiaomi open-sources MiMo-V2.5
    Xiaomi (MiMo)source ↗

    Alongside the Pro flagship, Xiaomi releases MiMo-V2.5, a ~310B/15B-active sparse-MoE model trained on ~48T tokens with a 1M-token context, under the MIT license.

  5. Released
    Xiaomi open-sources MiMo-V2.5-Pro
    Xiaomi (MiMo)source ↗

    Xiaomi releases its open-weight flagship MiMo-V2.5-Pro: a 1.02T-parameter MoE (~42B active) with hybrid attention and a 1M-token context, tuned for frontier-class agentic coding. MIT-licensed, with weights on Hugging Face.

  6. Released
    Tencent releases Hunyuan-A13B-Instruct
    Tencent Hunyuansource ↗

    80B/13B-active Hunyuan MoE model released with open weights and agentic tool-use support.

  7. Released
    OpenAI introduces GPT-Rosalind for life sciences
    OpenAIsource ↗

    OpenAI launches GPT-Rosalind, a frontier reasoning model purpose-built for biology, drug discovery, and translational medicine, available as a research preview in ChatGPT, Codex, and the API through a trusted-access program.

  8. Released
    Meta releases Muse Spark for Meta AI
    Meta AIsource ↗

    Meta releases Muse Spark, identified in reporting as the productized model behind Meta AI for U.S. users and the public successor to the Avocado effort.

  9. Updated
    Meta Avocado codename superseded by Muse Spark reporting
    Meta AIsource ↗

    The Avocado rumor row is retained for provenance; the public model is tracked separately as Muse Spark.

  10. Released
  11. Announced
    Anthropic discloses Claude Mythos but withholds public release
    Anthropicsource ↗

    Frontier model shipped only to ~50 defensive-security partners via Project Glasswing.

  12. Retired
    GPT-4o fully retired from ChatGPT
    OpenAIsource ↗

    Removed from all ChatGPT plans after a Feb 13 deprecation notice.

  13. Released
    Google DeepMind releases Gemma 4
    Google DeepMindsource ↗

    Gemma 4 introduces advanced reasoning open models in 12B, 26B, and 31B sizes.

  14. Released
    Z.ai launches GLM-5V-Turbo multimodal vision model
    Z.ai (Zhipu AI)source ↗

    Z.ai releases GLM-5V-Turbo, its first natively multimodal vision agent: image/video/text input with agent-oriented output (tool calling, task decomposition, GUI interaction) and a ~203K-token context.

March 2026

  1. Released
    Kimi K2.6 released
    Moonshot AIsource ↗
  2. Released
  3. Released
  4. Released
  5. Released
  6. Updated
    Reports say Meta delayed its Avocado model
    Meta AIsource ↗

    Meta's internal model reportedly missed a March target after underperforming top frontier systems; release timing and branding remain unconfirmed.

  7. Released
    Sarvam AI open-sources Sarvam-105B
    Sarvam AIsource ↗

    Apache-2.0 MoE model focused on reasoning, coding, agentic tasks, and Indian-language performance.

  8. Released
  9. Released
  10. Released
    Alibaba releases the small Qwen3.5 family (0.8B-9B)
    Alibaba (Qwen)source ↗

    Alibaba expands Qwen3.5 with four small dense models (9B, 4B, 2B, 0.8B) under Apache-2.0. The 9B is rated the most intelligent model under 10B parameters, with native vision and a 262K-token context.

  11. Released
  12. Released

February 2026

  1. Released
  2. Released
  3. Released
  4. Released
    OpenAI releases GPT-5.3-Codex
    OpenAIsource ↗

    OpenAI updates Codex with GPT-5.3-Codex, improving code quality, repository-scale reasoning, and long-running agentic coding workflows.

  5. Released
  6. Released

January 2026

  1. Released
    Moonshot releases Kimi K2.5
    Moonshot AIsource ↗

    Open multimodal K2 upgrade with MoonViT, thinking modes, visual coding, and agent-swarm workflows.

  2. Released

December 2025

  1. Released
    OpenAI releases GPT-5.2-Codex
    OpenAIsource ↗

    OpenAI releases GPT-5.2-Codex for software engineering agents, with a 400K-token input context and 128K-token output limit.

  2. Released
  3. Released
  4. Released
    OpenAI releases GPT-5.2
    OpenAIsource ↗

    OpenAI releases GPT-5.2 as a stronger general GPT-5 model for reasoning, coding, vision, instruction following, and long-context analysis.

  5. Released
  6. Released
  7. Released
  8. Released

November 2025

  1. Released
  2. Released
    Moonshot releases Kimi K2 Thinking
    Moonshot AIsource ↗

    Open K2 reasoning-agent variant for deep thinking and stable long-horizon tool orchestration.

October 2025

  1. Released
    Moonshot releases Kimi Linear 48B-A3B
    Moonshot AIsource ↗

    MIT-licensed hybrid linear-attention checkpoints with a 1M-token context and lower KV-cache use.

  2. Released
    Anthropic releases Claude Haiku 4.5
    Anthropicsource ↗

    Anthropic releases Claude Haiku 4.5 as a faster, lower-cost Claude 4.5 tier for coding, tool-use, and latency-sensitive agents.

September 2025

  1. Released
  2. Released
    Anthropic releases Claude Sonnet 4.5
    Anthropicsource ↗

    Anthropic releases Claude Sonnet 4.5, positioning it as its strongest model for coding, agents, and computer-use workflows at launch.

  3. Released
  4. Updated
  5. Released
    Moonshot updates Kimi K2 Instruct with 256K context
    Moonshot AIsource ↗

    The 0905 update improves agentic coding and frontend generation while doubling context length.

  6. Released

August 2025

  1. Released
  2. Released
  3. Updated
    Reports cite DeepSeek R2 delay tied to training hardware
    DeepSeeksource ↗

    R2 reportedly switched back to Nvidia for training after Huawei Ascend issues, with domestic hardware still targeted for inference.

  4. Released
  5. Updated
    Elon Musk says Grok 5 is planned before year-end
    xAIsource ↗

    The statement is tracked as a rumor until a public Grok 5 model release is confirmed.

  6. Released
    Anthropic releases Claude Opus 4.1
    Anthropicsource ↗

    Anthropic releases Claude Opus 4.1 with improvements for coding, reasoning, and agentic reliability over Claude Opus 4.

  7. Released
  8. Released
  9. Released
    Google releases Gemini 2.5 Deep Think
    Google DeepMindsource ↗

    Google makes its more deliberative Gemini 2.5 Deep Think reasoning mode available after previewing it at Google I/O 2025.

July 2025

  1. Released
    TII releases Falcon-H1 hybrid attention-SSM models
    Technology Innovation Institutesource ↗
  2. Released
  3. Released
  4. Released
    Qwen releases Qwen3-Coder-480B-A35B-Instruct
    Alibaba (Qwen)source ↗

    Alibaba Qwen releases its large open coding-agent MoE with 480B total / 35B active parameters and a 256K-token native context.

  5. Released
    Google releases Gemini 2.5 Flash-Lite
    Google DeepMindsource ↗

    Google releases Gemini 2.5 Flash-Lite as the lowest-cost, lowest-latency Gemini 2.5 tier for high-volume production tasks.

  6. Released
  7. Released
    Moonshot releases Kimi K2 Instruct
    Moonshot AIsource ↗

    Original open 1T-parameter K2 MoE release optimized for coding, reasoning, and agentic tool use.

  8. Released
  9. Released

June 2025

  1. Released
  2. Released
  3. Released
    Moonshot releases Kimi-VL-A3B-Thinking-2506
    Moonshot AIsource ↗

    Updated efficient multimodal reasoning model with stronger video, high-resolution perception, and lower thinking-token use.

  4. Released
    Google releases Gemini 2.5 Flash
    Google DeepMindsource ↗

    Google brings Gemini 2.5 reasoning improvements to a faster, lower-cost Flash production tier with a 1M-token context.

  5. Released
    Moonshot releases Kimi-Dev-72B
    Moonshot AIsource ↗

    Open coding LLM trained with repository-level reinforcement learning for issue resolution.

  6. Released
  7. Released
  8. Released

May 2025

  1. Released
  2. Released
  3. Released
  4. Preview
    Mistral previews Devstral Small 2505
    Mistral AIsource ↗

    Mistral and All Hands AI release Devstral Small 2505, a 24B Apache-2.0 coding-agent model for repository-level software engineering tasks.

  5. Released
  6. Released
    Google releases Gemma 3n E4B
    Google DeepMindsource ↗

    Google releases the mobile-first Gemma 3n E4B variant for efficient on-device multimodal inference.

  7. Released
    Mistral releases Mistral Medium 3
    Mistral AIsource ↗

    Mistral releases Medium 3 as a lower-cost enterprise workhorse for coding, STEM, search, and multilingual workloads.

April 2025

  1. Released
  2. Released
  3. Released
  4. Released
    Moonshot releases Kimi-Audio-7B-Instruct
    Moonshot AIsource ↗

    Open audio foundation model for speech recognition, audio QA, captioning, generation, and conversation.

  5. Released
    Moonshot releases Kimi-VL-A3B-Instruct
    Moonshot AIsource ↗

    Efficient MIT-licensed vision-language MoE for OCR, video, long documents, and agent tasks.

  6. Released
  7. Released
  8. Announced
    Meta previews Llama 4 Behemoth but does not release weights
    Meta AIsource ↗

    Meta described Behemoth as a still-training 288B-active / nearly 2T-total teacher model used to distill Llama 4 Scout and Maverick.

  9. Released
  10. Released
  11. Released

March 2025

  1. Released
  2. Released
  3. Released
  4. Released
  5. Released
  6. Released
  7. Released

February 2025

  1. Released
  2. Released
    Moonshot releases Moonlight-16B-A3B-Instruct
    Moonshot AIsource ↗

    Open 16B/3B-active MoE demonstrating Moonshot's scalable Muon optimizer work.

  3. Released
  4. Released
  5. Released

January 2025

  1. Released
  2. Released
    Qwen releases Qwen2.5-Max
    Alibaba (Qwen)source ↗
  3. Released
  4. Released
  5. Released
    Moonshot releases Kimi k1.5
    Moonshot AIsource ↗

    Multimodal reinforcement-learning reasoning model reported to match OpenAI o1 on math, coding, and multimodal reasoning.

  6. Released
  7. Released

December 2024

  1. Released
  2. Released
  3. Released
  4. Released
    TII releases the Falcon 3 small-model family
    Technology Innovation Institutesource ↗
  5. Released
  6. Released
  7. Released
  8. Released
  9. Released
  10. Released
  11. Released
  12. Released

November 2024

  1. Released
  2. Released
    Ai2 releases Tulu 3 405B
    Allen Institute for AI (Ai2)source ↗
  3. Released
  4. Released
  5. Released
  6. Released

October 2024

  1. Released
  2. Released
  3. Released
  4. Released
  5. Released
  6. Released

September 2024

  1. Released
  2. Released
    Ai2 releases Molmo 72B
    Allen Institute for AI (Ai2)source ↗
  3. Released
    Qwen releases Qwen2.5-72B
    Alibaba (Qwen)source ↗
  4. Released
  5. Released
  6. Released
  7. Released
  8. Released
  9. Released
    Ai2 releases OLMoE 1B-7B
    Allen Institute for AI (Ai2)source ↗

August 2024

  1. Released
  2. Released
  3. Released
  4. Released
  5. Released
  6. Released

July 2024

  1. Released
  2. Released

June 2024

  1. Released
  2. Released
  3. Released
  4. Released
  5. Released
    Qwen releases Qwen2-72B
    Alibaba (Qwen)source ↗
  6. Released
    Zhipu AI releases GLM-4-9B
    Z.ai (Zhipu AI)source ↗

May 2024

  1. Released
  2. Released
  3. Released
  4. Released
  5. Released
    TII releases Falcon 2 11B
    Technology Innovation Institutesource ↗
  6. Released
  7. Released

April 2024

  1. Released
  2. Released
    Snowflake releases Snowflake Arctic
    Snowflake AI Researchsource ↗
  3. Released
  4. Released
  5. Released
  6. Released
  7. Released
  8. Released
  9. Released
  10. Released

March 2024

  1. Released
  2. Released
  3. Released
    Databricks releases DBRX Instruct
    Databricks / MosaicMLsource ↗
  4. Released
  5. Released
  6. Released
  7. Released

February 2024

  1. Released
  2. Released
  3. Released
  4. Released
  5. Released
    Qwen releases Qwen1.5-110B
    Alibaba (Qwen)source ↗
  6. Released
    Qwen releases Qwen1.5-72B-Chat
    Alibaba (Qwen)source ↗

    Qwen releases the 72B chat-tuned Qwen1.5 checkpoint with 32K context and improved alignment.

  7. Released
    Ai2 releases OLMo 7B
    Allen Institute for AI (Ai2)source ↗

January 2024

  1. Released
  2. Released
    Zhipu AI releases GLM-4
    Z.ai (Zhipu AI)source ↗
  3. Released
  4. Released
  5. Released
  6. Released

December 2023

  1. Released
  2. Released
  3. Released
  4. Released
    Google announces Gemini 1.0
    Google DeepMindsource ↗

November 2023

  1. Released
  2. Released
  3. Released
    01.AI releases Yi-34B-Chat
    01.AIsource ↗

    01.AI releases the chat-tuned Yi-34B checkpoint alongside quantized chat variants.

  4. Released
  5. Released
  6. Released
  7. Released
  8. Released

October 2023

  1. Released
  2. Released

September 2023

  1. Released
  2. Released
  3. Released
  4. Released
    Qwen releases Qwen-14B
    Alibaba (Qwen)source ↗
  5. Released
  6. Released
  7. Released
    TII releases Falcon 180B
    Technology Innovation Institutesource ↗

August 2023

  1. Released
  2. Released
    Qwen releases Qwen-7B
    Alibaba (Qwen)source ↗

July 2023

  1. Released
  2. Released
  3. Released
  4. Released

June 2023

  1. Released
  2. Released

May 2023

  1. Released
    TII releases Falcon 40B
    Technology Innovation Institutesource ↗
  2. Released
  3. Released
    Databricks releases MPT-7B
    Databricks / MosaicMLsource ↗

March 2023

  1. Released
    LMSYS releases Vicuna 13B
    LMSYS / SkyLabsource ↗
  2. Released
  3. Released
  4. Released
  5. Released
  6. Released
  7. Released

February 2023

  1. Released

November 2022

  1. Updated
    ChatGPT (GPT-3.5) launches and reaches 100M users
    OpenAIsource ↗

    The consumer launch that ignited the modern LLM race.

  2. Withdrawn
    Meta pulls the Galactica demo after three days
    Meta AIsource ↗

    Withdrawn following criticism of confident but inaccurate scientific output.

  3. Released

July 2022

  1. Released

April 2022

  1. Released

December 2021

  1. Released
  2. Released

August 2021

  1. Released

June 2020

  1. Released

November 2019

  1. Released
    OpenAI releases the full GPT-2 (1.5B) weights
    OpenAIsource ↗

    After a staged rollout that began in Feb 2019 over misuse concerns.

October 2018

  1. Released
    Google releases BERT
    Google DeepMindsource ↗