LLM Releases
← Catalog

Nemotron 3.5 Lightning

Available
NVIDIAOpen weights

NVIDIA's efficient open Mixture-of-Experts model, released Aug 11 2026 for long-running agents. A hybrid Mamba-2 + MoE + Attention design with ~31.6B total and ~3.6B active parameters and a 1M-token context, shipped alongside the NeMo Switchyard model router. NVIDIA reports performance comparable to gpt-oss-120b at roughly a quarter of the total parameters, up to 4x the output speed of similar-sized models, and 10,000 tasks completed ~30% faster than Qwen3.6-35B at similar accuracy. Vendor-reported BF16 figures: SWE-bench Verified 51.56, GPQA Diamond 75.44, MMLU Pro 81.94, PinchBench 85.37 β€” self-reported and unverified by an independent harness at launch. Ships under the permissive OpenMDW-1.1 license with weights, training data, and recipes released, and is available on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com as an NVIDIA NIM microservice.

Specifications

License
Open weights Β· OpenMDW-1.1
Weights
Downloadable
Architecture
hybrid
Parameters
31.6B Β· 3.6B active
Context window
1M tokens
Max output
β€”
Knowledge cutoff
β€”
Price (in / out, $/M)
β€”
Modalities
TextCode

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.