Model family timeline
Last updated Sep 4, 2026
Ling model releases
A source-backed timeline for the Ling model family, collecting release dates, labs, access details, context windows, and major lifecycle changes.
4 models
Ling-3.0-flash-Sante
AvailableA health- and medicine-domain-tuned variant of Ant Group inclusionAI's Ling-3.0-flash, launched September 4 2026 (model id inclusionai/ling-3.0-flash-sante). Same efficient sparse Mixture-of-Experts base — 124B total parameters, ~5.1B active per token — post-trained for medical knowledge reasoning, clinical safety, evidence-based retrieval, and long-horizon medical tasks, while inclusionAI reports it retains strong general reasoning, coding, and agentic ability. Text-in / text-out only (no vision), with a 262,144-token (256K) context and up to 32,768 output tokens; supports reasoning and tool / function calling. Available first via hosted serverless API — Novita, OpenRouter (inclusionai/ling-3.0-flash-sante), and Vercel AI Gateway — with a time-limited free window at launch (free through Oct 4 on Vercel AI Gateway). Positioned as a developer API for research, retrieval, summarization, and workflow assistance, explicitly not a medical device or a substitute for clinical judgment; no public benchmark table at launch, so treat domain claims as unverified. Like the Fin variant, Sante-specific open weights were not confirmed posted at launch (the base Ling-3.0-flash family ships open-weight, MIT), so treat the open-weight status as announced-but-unconfirmed for this variant.
Ling-3.0-flash-Fin
AvailableA finance-domain-tuned variant of Ant Group inclusionAI's Ling-3.0-flash, launched August 27 2026. Same efficient sparse Mixture-of-Experts base — 124B total parameters, ~5.1B active per token — post-trained on high-quality financial data (developed with financial institutions and domain experts) for real-world investment and banking workflows: annual reports, financial workbooks, multi-document research, information retrieval, investment analysis, and valuation modeling, with an emphasis on complex multi-step tasks and long-horizon planning. Retains a 262,144-token (256K) context and up to 32,768 output tokens, and supports tool / function calling (tools and tool_choice), though it does not enforce structured JSON output (no response_format). inclusionAI reports it preserves strong general reasoning, coding, and math ability alongside the finance gains, citing finance benchmarks such as FinFIRST, FinSearchComp, and SpreadsheetBench — vendor claims, unverified independently. Launched hosted-API-first with a one-month free window through OpenRouter (inclusionai/ling-3.0-flash-fin); the promised open weights followed as announced and are now on Hugging Face under MIT (inclusionAI/Ling-3.0-flash-Fin, confirmed posted by 2026-09-04, with third-party hosting on DeepInfra and community GGUF quantizations).
Ling-3.0-tiny
AvailableThe smallest member of Ant Group inclusionAI's Ling 3.0 family, open-weighted on Hugging Face on Aug 6 2026 under the MIT license — distinct from the (API-only at launch) Ling-3.0-flash. A sparse Mixture-of-Experts model with 7.9B total parameters and only ~1.3B active per token: 128 routed experts with 8 routed plus 1 shared expert active per token, using the same 3:1 alternating stack of Kimi Delta Attention (KDA, linear) and Multi-head Latent Attention (MLA) layers as the rest of the family, for a 262,144-token (256K) context. Pitched as a highly economical on-device agent/reasoning model; weights are provided in BF16, FP8, and INT4 for a wide range of hardware. Vendor benchmark figures are unverified at launch.
Ling-3.0-flash
AvailableAnt Group's efficiency-focused Mixture-of-Experts model, released July 23 2026 by its inclusionAI lab: 124B total parameters activating only ~5.1B per token (1/64 expert activation). Ant claims it matches or beats the company's own ~1T-parameter Ling-2.6 flagship on most benchmarks it shows, at 1/8 the total and 1/12 the active parameters — a vendor claim with no public benchmark table or independent audit at launch, so treat it as unverified. Built for production-scale agents (MCP tool use, multi-agent coordination) rather than chat, with both thinking and non-thinking modes. Architecture is a native hybrid-linear attention stack interleaving Kimi Delta Attention (KDA) and Multi-head Latent Attention (MLA) at a reported 5:1 ratio, giving an economical 262,144-token (256K) context, with 1M cited as the scaling target. Ant docs claim peak inference up to 1,000 tokens/s and <100ms time-to-first-token on its own stack. Announced as open-weight under Apache 2.0, but as of July 24 no weights or model card were posted to the inclusionAI Hugging Face org — so the license and open-weight status are announced but unconfirmed (weights not yet downloadable; not self-hostable today). Usable now only via hosted API — free on OpenRouter (as inclusionai/ling-3.0-flash:free, hosted by Novita) and Vercel's AI Gateway through August 3 2026; no post-promo per-token price published at launch.