Latest LLM releases

LLM Releases
  1. DeepSeek promotes V4-Flash-0731 to production
    released - DeepSeek source

    DeepSeek ships DeepSeek-V4-Flash-0731, the production build of V4-Flash: the April preview retrained on an improved post-training pipeline for coding, agents, reasoning, and tool use, with the architecture untouched (284B total / 13B active MoE, 1M context). DeepSeek reports it beating its own larger V4-Pro-Preview on all nine published agent and coding benchmarks. Weights on Hugging Face under MIT; API pricing held at $0.14/$0.28 per Mtok, and the swap is silent for existing deepseek-v4-flash callers.

  2. V4-Flash API upgraded to the 0731 build
    updated - DeepSeek source

    The deepseek-v4-flash API endpoint is silently upgraded to the retrained DeepSeek-V4-Flash-0731 production build. Same endpoint, key, and model name; pricing unchanged at $0.14/$0.28 per Mtok.

  3. Alibaba quietly launches Qwen3.7-Flash
    released - Alibaba (Qwen) source

    Alibaba adds Qwen3.7-Flash, a cost-optimized vision-language reasoning model in the Qwen3.7 line, listed on OpenRouter and API endpoints on July 27 without a flagship announcement, technical report, or benchmark suite. It carries a 1M-token context and up to 65,536 output tokens, and at $0.03 / $0.13 per 1M input/output tokens is the cheapest 1M-context multimodal model available at release — aimed at high-volume multimodal agent workloads (visual coding, screen perception, browser/computer use, search) where cost matters more than peak intelligence. Closed-weights and API-only; architecture is undisclosed, with community speculation pointing to a small sparse-MoE design.

  4. Moonshot publishes Kimi K3 open weights
    updated - Moonshot AI source

    Moonshot AI publishes the full Kimi K3 weights to Hugging Face under a Modified MIT license on July 26 — a day ahead of its announced July 27 target — making the 2.8T-parameter MoE freely downloadable, modifiable, and self-hostable, and cementing K3 as the largest open-weight model publicly available.

  5. Anthropic releases Claude Opus 5
    released - Anthropic source

    Anthropic releases Claude Opus 5, its flagship model for demanding reasoning, autonomous coding, and long-horizon agentic work — pitched as the go-to model for most knowledge work, approaching Fable 5 capability in many categories at about half the price. Adds a five-level 'effort' dial on the Claude API/Platform to trade compute for capability, a 1M-token context at standard pricing, and up to 128K output tokens. Standard pricing $5/$25 per Mtok (matching Opus 4.8), plus a $10/$50 Fast mode; becomes the default for Claude Max subscribers. Anthropic calls it its most aligned Opus model.

  6. Ant Group's inclusionAI releases Ling-3.0-flash
    released - Ant Group (inclusionAI) source

    Ant Group's inclusionAI lab releases Ling-3.0-flash, a hybrid-reasoning Mixture-of-Experts model with 124B total parameters and ~5.1B active per token (1/64 expert activation), built for production-scale agents. Ant claims it matches or beats its own ~1T-parameter flagship on most benchmarks shown at 1/8 the total and 1/12 the active parameters — a vendor claim with no public benchmark table at launch. Uses a KDA + MLA hybrid-linear attention stack at a reported 5:1 ratio for an economical 256K-token context. Announced as open-weight under Apache 2.0, but weights and a model card were not yet posted to Hugging Face as of July 24; usable only via hosted API, free on OpenRouter and Vercel AI Gateway through August 3 2026.

  7. Google previews Gemini 3.5 Flash Cyber in CodeMender
    preview - Google DeepMind source

    A cyber-specialized model built on 3.5 Flash for finding and fixing software vulnerabilities, deployed inside the CodeMender agent and reaching competitive frontier performance on CyberGym. Limited-access pilot restricted to governments and trusted partners.

  8. Google releases Gemini 3.5 Flash-Lite
    released - Google DeepMind source

    Fastest, most cost-effective 3.5-class model (~350 output tokens/s) for high-throughput agentic workloads, priced at $0.30 / $2.50 per Mtok with a 1M-token context and built-in computer use. Large step up on 3.1 Flash-Lite and beats 3 Flash on several agentic/coding evals.