Ornith-1.5-397B
AvailableThe flagship of DeepReinforce's Ornith-1.5 family, released 2026-08-19 under the MIT license with weights on Hugging Face. A ~397B-parameter Mixture-of-Experts coding/agent model (per-token active count not disclosed) trained with a self-improving RL loop: rather than fixed human-curated tasks, the system proposes progressively harder tasks itself, generates a task-specific orchestration scaffold for each, and produces the solution rollouts used for reinforcement learning, with reward propagating across all three stages (all optimized with GRPO). Vendor-reported, five-run-averaged figures: Terminal-Bench 2.1 85.1 and DeepSWE 56.0 — which DeepReinforce puts on par with Claude Opus 4.8 (85.0 / 59.0) and ahead of GLM-5.2 and DeepSeek-V4-Flash-0731 at comparable scale — plus 92.8 GPQA Diamond and 86.6 BrowseComp. All numbers are the vendor's own and unverified by an independent harness at launch. Extends the self-scaffolding approach introduced in Ornith-1.0 (June 2026).
Specifications
- License
- Open weights · MIT
- Weights
- Downloadable
- Architecture
- Mixture-of-Experts
- Parameters
- Undisclosed
- Context window
- — tokens
- Max output
- —
- Knowledge cutoff
- —
- Price (in / out, $/M)
- —
- Modalities
- TextCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.