Xiaomi MiMo-V2.6: Live-streamed RL training that propels an open-source model to #1 on Artificial Analysis
🔎 Post-training becomes a public spectacle
On September 21, 2026, Xiaomi open-sourced the MiMo-V2.6 series under the MIT license: a giant 1.02-trillion-parameter model (MiMo-V2.6-Pro), a 309-billion Flash variant, and a distilled 9B checkpoint — all with weights on Hugging Face, the technical report and, the detail that changes everything, RL training logs viewable online.
Because MiMo-V2.6-Pro posts 46.32 on the Intelligence Index v4.3 from Artificial Analysis — the best score ever measured on an open-weights model under the v4.3 methodology, ahead of Kimi K3 and Qwen3.8 Max. And above all: Xiaomi didn't just publish the result. It published the recipe, the environments, and the cost.
Because the real story isn't the leaderboard. It's the reinforcement learning run that Xiaomi made public: 30 RL steps over roughly 750,000 trajectories, wrapped up in under six days, for $2.62 million in tokens on the Pro side and $854,000 on the Flash side. Precise, dated, verifiable figures — not a vendor benchmark. China is no longer content just catching up with the closed labs: it's industrializing reproducible post-training RL, and it's doing it in the open.
The Essentials
- MiMo-V2.6-Pro: 1.02 T parameters (42 B active, sparse MoE with 384 experts), omnimodal (text, image, audio, video input), 1M token context, MIT license, score of 46.32 on the Artificial Analysis Intelligence Index v4.3.
- MiMo-V2.6-Flash: 309 B parameters (15 B active), same omnimodal capabilities, designed to be served on reasonable hardware — the model a small team can both afford on the API and host themselves.
- Public RL run: 30 steps, 1,568 prompts × 16 rollouts per step, ~753,000 trajectories, 3.43–3.7 B tokens per step, stopped after 5 days and 7 hours (Pro) and 3 days and 11 hours (Flash), with costs detailed down to the dollar.
- 7,000+ open-sourced RL environments: software engineering, vulnerability reproduction, knowledge-intensive work, web development — with the complete framework to reproduce the training.
- API prices unchanged vs V2.5: $0.435/M input and $0.87/M output for Pro; $0.14/M and $0.28/M for Flash (September 2026, check openrouter.com).
- A Pro-UltraSpeed variant (up to 20x faster according to Xiaomi) rounds out the lineup on the API side.
Recommended Tools
| Tool | Main use | Price (September 2026) | Ideal for |
|---|---|---|---|
| MiMo AI Studio / Official API | Access to Pro, Flash and Pro-UltraSpeed | $0.435/M in / $0.87/M out (Pro); $0.14/M / $0.28/M (Flash) | Production, long-running agents (1M context) |
| MiMo-V2.6 weights on Hugging Face | Downloading the MIT weights, model card, benchmarks | Free | Research, fine-tuning, auditing |
| Live RL page | Viewing training logs (steps, pass rates, costs) | Free | Understanding an RL recipe at scale |
| OpenRouter | Unified Pro/Flash access with caching | Same pricing, cache hits at a fraction of a cent | Comparative multi-model testing |
| Ollama / LM Studio | Serving MiMo-V2.6-Distill-Qwen-9B locally | Free | Fine-tuning experimentation on consumer GPUs |
Three models, an MIT license, zero restrictions — what Xiaomi actually released
Direct answer: Xiaomi has released the full weights of three models under an MIT license, with no gating or usage clauses, making this the most permissive open-weights release at this level of performance to date.
The heart of the release is MiMo-V2.6-Pro. A sparse mixture-of-experts with 1.02 trillion parameters and only 42 billion activated per token, routed across 384 experts (8 active). Omnimodality is native: a 681M-parameter MiMo ViT vision encoder, a 308M AudioTokenizer plus a 127M patch encoder, all fused into a single model that outputs text. Xiaomi adds multi-token prediction with a 5-layer speculative decoder to speed up inference.
The 1M-token context isn't a marketing gimmick: the model card explicitly positions it for entire code repositories, tool traces, and multi-session agent runs. This is a model designed for agents, not for chatbots.
MiMo-V2.6-Flash (309B total, 15B active) is the pragmatic variant. According to Datanorth's analysis, it's "the model a four-person team can both afford on the API and serve on their own hardware." Pro, for its part, is for most organizations either a research asset or an inference bill — 1.02T of MIT-licensed weights won't fit on your workstation.
Finally, MiMo-V2.6-Distill-Qwen-9B: a checkpoint distilled from Qwen, released alongside Xiaomi's RL research resources. It's the model nobody mentions in the headlines, yet it will be the most downloaded — more on that below.
The live RL run: 30 steps, 750,000 trajectories, $2.6M — what the logs really show
Direct answer: for the first time, a lab has published the complete logs of an at-scale RL run — dates, costs, step-by-step pass rates — and it looks as much like a deliberate transparency exercise as a show of infrastructure muscle.
The numbers, sourced from the official RL page:
| Metric | MiMo-V2.6-Pro | MiMo-V2.6-Flash |
|---|---|---|
| Run start | Sept 15, 2026, 10:32 UTC | Sept 15, 2026, 15:16 UTC |
| Run end | Sept 20, 2026 (step 30) | Sept 19, 2026 (step 30) |
| Duration | 5 d 07 h 29 min | 3 d 11 h 05 min |
| Final dynsam/avg@n score | 0.633 (+0.068 vs step 1) | 0.644 (+0.130 vs step 1) |
| Relative gain | ~+12% | ~+25% |
| Cumulative token cost | $2,620,670 | $854,044 |
| Tokens at step 30 | 3.43B | 3.7B |
| Training trajectories | ~753,000 | ~753,000 |
| Batch | 1,568 × 16 sequences | 1,568 × 16 sequences |
Two things deserve your attention. First, the fully asynchronous architecture: each update consumes 1,568 samples and trains on 3.5 to 3.7 billion tokens with a 1M context. This isn't LoRA fine-tuning on a single GPU — it's an industrial pipeline of rollout generation, scoring, and updates, sustained under load for several days.
Next, the method. The paper submitted on September 21 describes an approach based on comparing "sibling" attempts within each batch: rather than scoring each rollout in absolute terms, the model learns by comparing attempts against one another to recenter the signal toward the cleanest solution. This graded feedback is applied simultaneously across four domains — code, professional workflows, visual design, and cybersecurity — and the score-vs-step curves show steady progress on all four.
The most telling result remains DeepSWE v1.1: the run takes Flash from 19.0 (V2.5) to 65.68 on the mini-swe-agent avg@3 evaluation from the RL page — a jump of nearly 47 points, a large share of which was gained during these 30 steps. This is where post-training RL stops being an alchemical art and becomes a measurable process.
My take: the most interesting thing isn't the final score, it's the cost per point of gain. ~$2.6M for +12% relative on Pro, ~$0.85M for +25% relative on Flash. For the first time, a team considering a post-training RL run has public orders of magnitude to budget against. It may not look like much; it's going to shorten decision cycles for quite a few mid-sized labs.
46.32 on the v4.3 Index: how does this score stack up against Kimi K3 and Qwen3.8?
Direct answer: on Artificial Analysis's v4.3 Index, MiMo-V2.6-Pro is the highest-ranked open weights model — but the numbers need to be read with the rigor that comparing across index versions demands.
According to Artificial Analysis, MiMo-V2.6-Pro scores 46 on the Intelligence Index, "well above the average of other similarly sized open weights models (median: 18)". Datanorth puts it at 46.32 on v4.3, tied with Grok 4.7 — released the same day by SpaceXAI, and roughly five times more expensive per token. Grok 4.6 sits at 44, Gemini 3.8 Flash at 41.
An honest point of caution: the comparisons with Kimi K3 and Qwen3.8 that are circulating use figures from an earlier version of the Index, where K3 was credited with 60 and Qwen3.8 with 58. These figures, taken from a vendor article by Featherless, are not directly comparable to v4.3. What remains solid are the cross-benchmarks on agentic tasks, where MiMo-V2.6-Pro plays in the same league as the closed models:
| Benchmark | MiMo-V2.6-Pro | MiMo-V2.6-Flash | Closed reference |
|---|---|---|---|
| DeepSWE v1.1 | 71.9 | 67.9 | Claude Opus 5: 74.0 |
| Terminal Bench 2.1 | 89.9 | 87.6 | — |
| OSWorld-Verified | 82.0 | 80.8 | — |
| Toolathlon-Verified | 76.9 | 73.6 | — |
| AutomationBench v1.0.6 | 53.1 | 52.3 | — |
| Agents' Last Exam | 31.6 | 27.6 | — |
| ExploitBench | 47.9 | 25.3 | — |
Goldie Bench estimates that Pro costs about a quarter of Grok 4.7's price and one-twentieth that of closed frontier models, for an equivalent level on agentic benchmarks. On DeepSWE, the gap with Claude Opus 5 (74.0) is now down to just 2.1 points. In one year, open weights went from "interesting but 15 points behind" to "right on their heels".
Xiaomi's official doc also claims the integration of 3D spatial reasoning, multimodal perception, and a Computer Use Agent — understanding of complex graphical interfaces, manipulation of office tools, verification of results against visual feedback — plus a "Vibe World" for interactive world-building driven by natural language. The OSWorld-Verified score of 82.0 suggests this isn't just hot air.
7,000 open RL environments: the real bombshell of the release
Direct answer: the model will grab the headlines; it's the 7,000+ RL task environments and the reproduction framework that will change post-training practice in the open source community.
The official doc describes "7k+ high-quality RL task environments covering four types of agentic tasks: software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development." Xiaomi is simultaneously publishing the verified training practices and the complete framework to relaunch the run.
Why does this matter? Because the bottleneck of post-training RL has never been the algorithm — PPO, GRPO and variants have been known for years. It's the quality of the environments: tasks with reliable verifiers, calibrated difficulty, and enough diversity to avoid reward hacking. Building 7,000 environments of this kind (including vulnerability reproduction, which is rare and sensitive) probably represents more engineering than the run itself.
By open-sourcing this corpus under MIT, Xiaomi is doing something rather subversive: it's turning a skill held by only five labs in the world into reproducible infrastructure. Any team with the GPU budget can now start from a documented recipe rather than reinventing its pipeline. This is exactly the kind of move that, historically, accelerates an entire category of models — and that gives meaning to LLM for agents comparisons beyond simple chat scores.
Want to fine-tune an agent? Start with the Distill-Qwen-9B
Direct answer: nobody fine-tunes 1.02T parameters. The realistic experimentation ground is MiMo-V2.6-Distill-Qwen-9B — and that's precisely why Xiaomi is releasing it along with its RL resources.
The pipeline logic is explicit in the official doc: RL training starts from the distilled checkpoint. Xiaomi therefore designed this 9B as a viable starting point for reinforcement learning, not as a mere demo model. Concretely:
- Download the weights from Hugging Face (MIT license, commercial use included).
- Serve it locally with Ollama or LM Studio — a 9B runs comfortably on a 24 GB consumer GPU, and a small dedicated server is enough for an internal service; for that, a GPU VPS from Hostinger does the job without any hardware investment.
- Adapt Xiaomi's RL environments to your business domain: replace their web dev tasks with your own verifiable tasks, and keep their scoring framework.
- Iterate by comparing siblings — the method from the paper is reproducible at small scale, and it's what stabilizes the learning signal.
For cases where the 9B isn't enough, Flash remains the intermediate option: 15B active parameters, DeepSWE at 67.9, and a cost profile that stays manageable. Our guide to the best local LLMs will be updated with these new benchmarks, and for choosing the base model of an agentic pipeline, the monthly comparison of the best LLMs remains the go-to reference.
The September 2026 open weights wave: post-training becomes a Chinese industry
Direct answer: MiMo-V2.6 doesn't land in a vacuum. It arrives in a September 2026 where China has flooded the open weights ecosystem with coordinated releases — and the strategy is now legible.
The month's roundup:
- DeepSeek V4-Pro under MIT license, with the recent episode of V4-Pro being swapped for V4.1-Flash behind the same endpoint — detailed analysis in DeepSeek swaps V4-Pro for V4.1-Flash.
- Qwen3.8-2.4T-A95B, Alibaba's open checkpoint: 2.4T total parameters, 95B active, 262K native context extendable to ~1M, reasoning that cannot be disabled — and a 4.89 TB full-precision checkpoint, which gives an idea of the serving cost of this generation.
- Kimi K3 from Moonshot AI: a 2.8T MoE, 104B active, MXFP4 weights (~1.4 to 1.56 TB), native vision via MoonViT-V2.
- GLM-5.2 from Z.AI, presented as the most powerful open weights model in the world at its release (753B MoE, 1M context) — see GLM-5.2: the most powerful open weights model in the world.
- StepFun Step-5 Preview, a 600-billion-parameter MoE with 1M context (analysis here).
- Ternary-Bonsai-2 and its 2-bit approach, which pushes weight compression where no one expected it — tracking of these releases is handled by The Open Weights.
Faced with this, America is showing signs of strategic retreat. Meta has taken the plunge into closed models with Muse Spark — a betrayal of open source analyzed in Meta Muse Spark: why Meta betrayed open source — while NVIDIA counterattacks with Nemotron 3 Ultra 550B and Google explores another path with DiffusionGemma, the first open source text diffusion model.
My read: China has understood that the war over base models is shifting toward reproducible post-training. 1T weights can be copied; RL pipelines with 7,000 environments, verified training practices, and public costs — that's industrial capacity exported as code. Xiaomi, a maker of smartphones and agentic cars, applies to RL the same manufacturing logic it applies to its hardware — as mPost highlights on agentic cars, cyber, and 3D worlds.
❌ Common Mistakes
Mistake 1: comparing Index scores without checking the version
The figures "K3 at 60, Qwen3.8 at 58" have been circulating since an earlier version of the Artificial Analysis Index. MiMo-V2.6-Pro's 46.32 is measured on v4.3. Mixing the two is like comparing different thermometers. The solution: always check the Index version and the capture date on artificialanalysis.ai before drawing conclusions.
Mistake 2: believing you'll "run Pro at home"
1.02 T parameters, even with 42 B active, implies a multi-GPU infrastructure for which Xiaomi publishes no hardware recommendations anyway. The self-hostable model for an individual or a small team is Distill-Qwen-9B — or possibly Flash if you have serious hardware. Everyone else will go through the API.
Mistake 3: underestimating the cost of an RL run
The public logs give the real orders of magnitude: $854,000 in tokens for 30 steps on Flash, $2.62M on Pro, excluding infrastructure and engineering. Many teams budget for SFT fine-tuning and then discover that RL costs one to two orders of magnitude more. Start from Xiaomi's figures, not from intuition.
Mistake 4: confusing benchmark scores with a model that's useful in production
MiMo-V2.6-Pro is described as fast but "somewhat verbose" by Artificial Analysis. On Terminal Bench 4.0, it stays at 34.9 — the task is hard for everyone. A good overall agentic score tells you nothing about your specific use case: test on your own tasks before migrating.
❓ Frequently Asked Questions
Is MiMo-V2.6-Pro really usable commercially?
Yes. The MIT license covers all three models, with weights downloadable from Hugging Face and commercial use included, with no restrictions or non-compete clauses. It's the most permissive license on the market for this level of performance. The only practical constraint: the size of Pro's weights, which makes the API necessary for most use cases.
Which model should you choose between Pro, Flash, and the 9B?
Pro for maximum performance via API (long-running agents, 1M context, OSWorld 82.0). Flash if you want a servable compromise: 15 B active parameters, DeepSWE 67.9, at one-third the price. The distilled 9B for local experimentation and RL fine-tuning. Our comparison of LLMs for code details the trade-offs by use case.
Is the RL run really reproducible with the published materials?
Xiaomi publishes the framework, the 7,000+ environments, the verified training practices, and the hyperparameters visible on the RL page (batch 1,568 × 16, ~3.5 B tokens per step). Reproducing it exactly will cost millions of dollars, but reproducing the method at small scale on the 9B is realistic for a well-equipped team.
How do these prices compare to closed models?
Pro: $0.435/M input, $0.87/M output. Flash: $0.14/M and $0.28/M (September 2026, verify on openrouter.com). According to Goldie Bench, that's roughly a quarter of the price of Grok 4.7 and one-twentieth that of closed frontier models, at equivalent agentic benchmark levels. Cache hits drop to a fraction of a cent.
Why is publishing training costs so important?
Because it's a piece of data that almost no lab discloses. Knowing the real cost of an RL run ($2.6M for a relative +12% on Pro) lets any team decide whether to launch their own, and to negotiate their infrastructure with figures rather than rumors. It's a rare act of transparency — and a strategic one.
✅ Conclusion
MiMo-V2.6 is the release that takes open weights from "credible follower" status to that of measured leader on the Index v4.3 — but its true legacy is the public RL run, the roughly 7,000 environments, and the costs published down to the dollar, which turn post-training into a reproducible discipline. If you want to test what this changes for your agents, start with the local LLM installation guide and Distill-Qwen-9B: that's where the next generation of open source models will be written.