LLM Releases

Lab release history

Last updated Aug 19, 2026

DeepReinforce (Ornith) model releases

AI research startup founded by Jiwei Li (Stanford CS PhD, earlier founder of the NLP startup Shannon.AI); focuses on reinforcement-learning optimization for coding and agents, with prior open RL work including CUDA-L1 and the IterX agent loop. Ships the open-weight, MIT-licensed Ornith family, whose self-scaffolding / self-improving RL loop has the model propose its own tasks, generate an orchestration scaffold for each, and produce the training rollouts. Country of incorporation is not clearly documented in public sources — recorded as US (Stanford-affiliated founder) pending confirmation. This page collects the lab's model releases, lifecycle events, source links, and model metadata in one crawlable record.

3
Models
1
Labs
3
Open
3
Recent

3 models

Ornith-1.5-9B

Available
DeepReinforce (Ornith)Open weights

The smallest model in DeepReinforce's Ornith-1.5 family (released 2026-08-19, MIT, weights on Hugging Face): a 9B-parameter dense coding/agent model trained with the family's self-improving task-and-scaffold RL loop, and shipped with a quantized 'Ornith-1.5-9B-Mobile' build that runs on iPhone and Android. Vendor-reported, five-run-averaged figures put it at 47.0 on Terminal-Bench 2.1 and 70.6 on SWE-Bench Verified, which DeepReinforce places above larger models including Gemma 4-31B and Qwen3.6-35B-A3B. Figures are self-reported and unverified at launch.

Dense9B ctxAug 19, 2026

Ornith-1.5-35B-A3B

Available
DeepReinforce (Ornith)Open weights

The mid-size model in DeepReinforce's Ornith-1.5 family (released 2026-08-19, MIT, weights on Hugging Face): a 35B-parameter Mixture-of-Experts that activates ~3B parameters per token, trained with the same self-improving task-and-scaffold generation loop as the 397B flagship. Vendor-reported, five-run-averaged figures put it at 68.5 on Terminal-Bench 2.1 and 79.0 on SWE-Bench Verified while activating only 3B parameters per token — which DeepReinforce reports as outperforming dense models of similar or larger size such as Meta's Muse-Glimmer-30B and Gemma 4-31B. Figures are self-reported and unverified at launch.

MoE35B ctxAug 19, 2026

Ornith-1.5-397B

Available
DeepReinforce (Ornith)FrontierOpen weights

The flagship of DeepReinforce's Ornith-1.5 family, released 2026-08-19 under the MIT license with weights on Hugging Face. A ~397B-parameter Mixture-of-Experts coding/agent model (per-token active count not disclosed) trained with a self-improving RL loop: rather than fixed human-curated tasks, the system proposes progressively harder tasks itself, generates a task-specific orchestration scaffold for each, and produces the solution rollouts used for reinforcement learning, with reward propagating across all three stages (all optimized with GRPO). Vendor-reported, five-run-averaged figures: Terminal-Bench 2.1 85.1 and DeepSWE 56.0 — which DeepReinforce puts on par with Claude Opus 4.8 (85.0 / 59.0) and ahead of GLM-5.2 and DeepSeek-V4-Flash-0731 at comparable scale — plus 92.8 GPQA Diamond and 86.6 BrowseComp. All numbers are the vendor's own and unverified by an independent harness at launch. Extends the self-scaffolding approach introduced in Ornith-1.0 (June 2026).

MoEUndisc. ctxAug 19, 2026