Latest LLM releases

LLM Releases
  1. Salesforce and NVIDIA announce Koa, Salesforce's first CRM reasoning model
    released - Salesforce source

    At Dreamforce, Salesforce and NVIDIA announced Koa, Salesforce's first CRM reasoning model for Agentforce, built by post-training NVIDIA Nemotron 3 Super on a proprietary synthetic dataset modeled on ~27 years of CRM deployments (no customer data used). Koa reasons through multi-step enterprise workflows and uses tools to act; on Salesforce's CRM benchmark it matches or exceeds leading models on CRM actions with roughly 3x fewer errors. Salesforce controls the weights and runs inference inside its own trust boundary (weights not released). Available to select pilot customers at launch; general availability expected winter 2026 in U.S. regions.

  2. TypeSafe AI launches Jev, its first System One automation model
    released - TypeSafe AI source

    TypeSafe AI released Jev in early access — the first model in its "System One" line, built for automation workflows rather than chat. Jev maps unstructured state to typed probabilistic decisions and emits parallel structured outputs (JSON / tool calls), trained with Reinforcement Learning for Calibrated Decisions (RLCD), and is positioned as a "frontier-intelligence function call" for agents, classification, and tool use. Proprietary (conditional commercial use); weights not released; context window undisclosed. Priced at $0.042 /1M input with output free on a single TypeSafe AI serverless route at launch. Vendor figures unverified independently.

  3. Singapore's Agnes AI publishes the open-weights Agnes 3.0 Flash preview
    released - Agnes AI source

    Agnes AI, a Singapore omni-modal foundation-model lab, surfaced Agnes 3.0 Flash in mid-September 2026. The disclosed open-weights preview checkpoint (Apache 2.0) is a 33B hybrid-attention model with a 262K context and text/image/video input: 54 of 72 decoder layers run a gated delta rule while 18 use standard grouped-query attention, holding down the KV cache at long context. Runs at bf16 on a single H100/H200-class GPU. The production Agnes 3.0 Flash served via the company's API is a separate checkpoint with a 1M-token context. First model tracked from this org. Vendor figures unverified at launch.

  4. All V4-Pro API traffic routes to V4.1-Flash
    deprecated - DeepSeek source

    From 04:00 UTC on 2026-09-14 DeepSeek routes every `deepseek-v4-pro` request to V4.1-Flash, billed at V4.1-Flash rates, and says this will continue until V4.1-Pro launches. The id still answers, making this a redirect and a deprecation rather than a retirement — DeepSeek reports V4.1-Flash beats V4-Pro on performance, cost, speed, and total time.

  5. Inference.net lists the Schematron V2 extraction models
    released - Inference.net source

    Inference.net listed Schematron V2 Turbo and Small, a pair of 3B HTML-to-JSON extraction models in its "workhorse model" line — small purpose-built LLMs sold on cost per unit of work. Both are schema-driven (the extraction target goes in a JSON schema via response_format, not the prompt) with 128K context; Turbo is throughput-optimized at ~4.14 req/s on one H100 and $0.03/$0.15 per Mtok, Small trades throughput for quality on complex schemas at $0.05/$0.23. Proprietary and API-only via Inference.net and OpenRouter. First models tracked from this org.

  6. Shanghai AI Lab ships Atria Dawn Preview, a 744B open-weight agentic MoE
    released - Shanghai AI Laboratory source

    Shanghai AI Laboratory released Atria Dawn Preview weights-first on 2026-09-11 — code and checkpoint appeared on GitHub / Hugging Face with no announcement, followed ~three days later by a 140-author technical report. It is a 744B-parameter agentic Mixture-of-Experts model built on GLM-5.2, aimed at long-horizon research agents that take a method from the literature to executable experiments, reproducible metrics, and an inspectable report. MIT license, open weights, 256K context. On the lab's own 16-benchmark table it leads on five tasks incl. AutomationBench 53.8, BrowseComp 92.5, DeepSearchQA 96.0, BFCL v4 77.0 and CyberGym 86.5. Self-reported figures.

  7. Moonshot ships Kimi K2.8 Preview on Kimi Code
    released - Moonshot AI source

    Moonshot AI released Kimi K2.8 Preview, a mid-tier coding and agentic model positioned between Kimi K3 and Kimi K2.7 Code, with performance Moonshot describes as close to K3 but with significantly more efficient thinking. It brings the K3-series thinking-effort controls (low/high/max, max default), multimodal input (text, image, video) with text output, and a 1M-token context now available across all membership tiers. Served on Kimi Code under model id kimi-for-coding so existing clients pick it up without config changes. Proprietary, preview status; parameters undisclosed; vendor figures unverified at launch.

  8. Sakana AI ships Fugu Ultra v2 without frontier models in its pool
    released - Sakana AI source

    Sakana AI released Fugu Ultra v2.0, the second generation of its frontier-class orchestration model, reporting frontier-level results with no Claude Fable 5, Fable 5.1, or GPT-6 Astra among its agents — routing instead over open-weights and specialized models including NVIDIA Nemotron, which Sakana positions as resilience against single-vendor and export-control risk. Tuned for sustained reasoning over complex visual and structured data (Sakana-reported 48.3 Chartography, 74.3 DeepSWE). $5/$30 per Mtok standard, $10/$45 above 272K context. Figures describe an orchestrated system, not a single set of weights.