Lab release history
Last updated Sep 21, 2026
Xiaomi (MiMo) model releases
Chinese consumer-electronics giant; its MiMo team ships open-weight, MIT-licensed MoE reasoning and coding models. This page collects the lab's model releases, lifecycle events, source links, and model metadata in one crawlable record.
4 models
MiMo-V2.6-Pro
AvailableMiMo-V2.6-Pro is Xiaomi MiMo's flagship open-weights model released 2026-09-21, succeeding the proprietary MiMo-V2.5-Pro. Sparse Mixture-of-Experts (MiMoV2ForCausalLM): 1.02T total / 42B active parameters, 70 layers (60 sliding-window + 10 global attention), 384 routed experts with 8 activated per token, hybrid SWA+GA backbone, 681M MiMo ViT vision encoder, and MTP blocks. Natively omnimodal (text, image, video, audio input; text output); 1,048,576-token context window and 128,000 max output tokens. Open weights on Hugging Face (XiaomiMiMo/MiMo-V2.6-Pro-RL) under MIT, ungated. Trained with a "You Only RL Once" mixed-RL recipe across coding, agents, visual, and cybersecurity. Overseas API pricing: $0.435 / 1M input (cache-hit $0.0036), $0.87 / 1M output. Artificial Analysis reports an Intelligence Index of 46, a top open-weight score.
MiMo-V2.6-Flash
AvailableMiMo-V2.6-Flash is Xiaomi MiMo's cheaper open-weights sibling to MiMo-V2.6-Pro, released 2026-09-21. Mixture-of-Experts with 309B total / 15B active parameters; 48 layers (1 dense + 47 MoE), 256 routed experts (top-8), hybrid sliding-window (128) + global attention. Natively omnimodal (text, image, video, audio input; text output); 1,048,576-token context window and 128,000 max output tokens. Open weights on Hugging Face (XiaomiMiMo/MiMo-V2.6-Flash-RL) under MIT. Trained on ~48T tokens (26T text stage + 22T multimodal stage). Optimized for agentic workflows across coding, visual, general, and research scenarios; nearly matches Pro at about a third of the API price. Pricing: $0.14 / 1M input (cache read $0.0028), $0.28 / 1M output.
MiMo-V2.5-Pro
AvailableXiaomi's open-weight flagship: a 1.02T-parameter Mixture-of-Experts model with ~42B active parameters, a hybrid-attention architecture, and a 1M-token context window. Tuned for frontier-class agentic coding and long-horizon tasks (sustaining 1000+ tool calls with a proper harness). Open-sourced under the MIT license with weights and tokenizer on Hugging Face.
MiMo-V2.5
AvailableXiaomi's open-weight sparse-MoE model: ~310B total parameters with ~15B active, trained on ~48T tokens, with a 1M-token context window. Shipped alongside the larger MiMo-V2.5-Pro under the MIT license.