LLM Releases
← Catalog

Naive-N0.5-Flash

Available
NaiveAIOpen source

Naive-N0.5-Flash is the debut open-weights model from NaiveAI, a Beijing-based lab, released 2026-09-27 under the MIT license with weights and inference code published on Hugging Face. It is a 309B-total / ~15.5B-active Mixture-of-Experts model continually pre-trained on about 3.25T additional tokens from Xiaomi's open-weight MiMo-V2.5. Its headline feature is an attention stack with NO full/global-attention layers: 48 transformer layers made up of 39 sliding-window-attention layers (128-token window) and 9 DeepSeek Sparse Attention layers (which select the top ~2,048 most relevant tokens), with GQA-4 grouping — giving a native 1,048,576-token (1M) context at reduced memory and compute. Text input and text output, aimed at coding and AI research/engineering workloads, with an "Ultrafast" serving mode cited at up to ~2,000 tokens/sec. A hosted API launched invitation-only at $0.10 / 1M input and $0.40 / 1M output ($0.01 / 1M cached reads). NaiveAI reports results on coding benchmarks (SWE-Bench Pro, DeepSWE v1.1, Terminal-Bench 2.1) and AI-R&D benchmarks (PostTrainBench, MLE-bench-30, PaperBench), but publishes them only as figures/plots rather than a numeric table, and the model does not yet appear on the public SWE-bench Pro leaderboard. A derivative/continued-pretrain of MiMo-V2.5 rather than a from-scratch model; figures are vendor / self-reported.

Specifications

License
Open source · MIT
Weights
Downloadable
Architecture
Mixture-of-Experts
Parameters
309B · 15.5B active
Context window
1.0M tokens
Max output
—
Knowledge cutoff
—
Price (in / out, $/M)
$0.1 / $0.4
Modalities
TextCode

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.