Nemotron 3.5 Lightning
AvailableNVIDIA's efficient open Mixture-of-Experts model, released Aug 11 2026 for long-running agents. A hybrid Mamba-2 + MoE + Attention design with ~31.6B total and ~3.6B active parameters and a 1M-token context, shipped alongside the NeMo Switchyard model router. NVIDIA reports performance comparable to gpt-oss-120b at roughly a quarter of the total parameters, up to 4x the output speed of similar-sized models, and 10,000 tasks completed ~30% faster than Qwen3.6-35B at similar accuracy. Vendor-reported BF16 figures: SWE-bench Verified 51.56, GPQA Diamond 75.44, MMLU Pro 81.94, PinchBench 85.37 β self-reported and unverified by an independent harness at launch. Ships under the permissive OpenMDW-1.1 license with weights, training data, and recipes released, and is available on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com as an NVIDIA NIM microservice.
Specifications
- License
- Open weights Β· OpenMDW-1.1
- Weights
- Downloadable
- Architecture
- hybrid
- Parameters
- 31.6B Β· 3.6B active
- Context window
- 1M tokens
- Max output
- β
- Knowledge cutoff
- β
- Price (in / out, $/M)
- β
- Modalities
- TextCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.