Xiaomi MiMo-V2.6: The Open-Weight Model That Costs $0.13 a Task

A phone maker just shipped the highest-scoring open-weight model on Artificial Analysis, trained in under six days for $2.62M. Here's what MiMo-V2.6-Pro actually takes to run.

Cover art for Xiaomi MiMo-V2.6: The Open-Weight Model That Costs $0.13 a Task

Xiaomi - the phone and EV company - released an open-weight model on September 22 that scored higher on Artificial Analysis's Intelligence Index than any open-weight model before it. MiMo-V2.6-Pro is the highest-scoring open-weight model on Artificial Analysis's index, and Xiaomi says it trained the whole thing in under six days for about $2.62 million. That is not a typo, and the weights are on Hugging Face under MIT. The gap between that headline and what it actually takes to run the model is where this post lives.

46Artificial Analysis Intelligence Indexhighest ever for an open-weight model
$2.62Mreported training costsix days of compute
$0.13cost per AA Intelligence Index taskcheapest model Artificial Analysis tracks
1.02T / 42Bparameters total / active per tokenMoE, not a dense model

What MiMo-V2.6-Pro actually is

MiMo-V2.6-Pro is Xiaomi's open-weight flagship, released September 22 under MIT. It is a 1.02T-parameter MoE with 42 billion active parameters per forward pass, a 1M-token context window, and native text, image, video, and audio input. The mixture-of-experts design matters here: only 42B parameters fire for any given token, which is why the inference cost stays low even though the model is formally a trillion-parameter system.

The team was led by Luo Fuli, formerly of DeepSeek

  • a detail that explains a lot about the architecture decisions, since the MoE routing and training approach look very close to the DeepSeek V4 playbook.

Xiaomi published the training code and 7,000 task environments alongside the weights , which is unusual. Most labs drop weights and a technical report. Publishing the RL environments means the community can reproduce the post-training pipeline, not just serve the checkpoint.

On coding and agent benchmarks, Xiaomi's own table has Pro reaching 89.9% on Terminal Bench 2.1, above Claude Opus 5 at 89.1%, and 94.0% on CyberGym against 40.0% for MiMo-V2.5-Pro. The cybersecurity number is not a minor side stat - it's a 135-point jump from one generation to the next, which suggests the RL post-training included adversarial task environments that earlier versions were not trained on. Take vendor-published benchmarks at face value at your own risk, but independent evaluators land in roughly the same place: Artificial Analysis's composite index draws on 10 evaluations including AA-Briefcase v1.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam.

The jump from the previous generation is steep. V2.6 came five months after MiMo-V2.5, and on Artificial Analysis's Intelligence Index, V2.5-Pro scored 26 and V2.6-Pro scores 46. A 77% improvement on an independent composite in five months is either a genuine architectural leap, an unusually well-targeted RL training run, or both.

How to actually run it - and what the docs require

This is where the headline diverges from the workload. Running MiMo-V2.6-Pro locally means serving a multi-node GPU cluster with tensor and expert parallelism through SGLang or vLLM, not loading a checkpoint on a single workstation.

The vLLM recipe is the easier path. Hardware: 8× H200 (TP8), 4× MI355X (TP4), or equivalent aggregate VRAM of at least 680 GB.

Weights alone need about 72 GB per GPU, so in practice that means 8× H200 (141 GB each) or 8× B200-class cards. 8× H100 80 GB leaves almost no room for KV cache.

The SGLang production path is heavier. Xiaomi's tuned SGLang recipe uses two nodes - TP 16, DP 2, expert parallelism 16, DeepEP all-to-all, EAGLE multi-layer speculative decoding.

There is also a version mismatch gotcha. The vLLM recipe documents a 566 GB on-disk checkpoint and requires a special image or nightly vLLM build rather than the then-current stable release , because stable vLLM ≤ 0.29.0 cannot load the MXFP4-stored weights. If you reach for pip install vllm today, you will not be able to serve the model.

Flash is a more tractable entry point. The Flash checkpoint requires 4× H200 (TP4) or equivalent with at least 208 GB aggregate VRAM, and the weights are only 173 GB on disk thanks to MXFP4 storage with FP8 computation.

Variant AA Index Total params Active params Min VRAM Disk size License
Pro-RL 46 1.02T 42B 680 GB 566 GB MIT
Flash-RL ~38 ~200B ~20B 208 GB 173 GB MIT
Distill-Qwen-9B - 9B 9B (dense) ~20 GB ~18 GB MIT

The launch also includes Flash weights, a 9-billion-parameter distillation, a technical report, training environments, and reinforcement learning code

  • the 9B distillation is the one that runs on a single GPU and will probably generate the most real-world adoption.

What the pricing actually looks like

Pro's API costs $0.435 input and $0.87 output per million tokens through Xiaomi. Compare that to Kimi K3, which sits at 44 on the same index. Pro costs about one-seventh of Kimi K3 on input and one-seventeenth on output.

Cached input at $0.004 per million makes long-running agents with a stable system prompt and repository context close to free on the input side. For a coding agent that runs against the same 100K-token codebase all day, the effective per-call cost collapses.

An MIT license on a trillion-parameter omni-modal model is unusual - teams can self-host, fine-tune, and ship MiMo-V2.6 commercially with no usage restrictions. That is the part that matters for teams with data residency requirements or who are building a product on top of the model: no API terms, no usage clause, no key rotation.

Beagle in action#engineering, 11:22am
The ask
'can we route the agent loop through MiMo-V2.6-Flash instead of the API we're paying for now?'
Beagle drafts
pulls the OpenRouter listing and Xiaomi pricing, drafts a cost comparison against current spend with a note on the 208 GB VRAM requirement
You approve
you approve the reply; the team has the numbers before standup
Do this in your workspace →

The non-obvious thing: it's a phone maker, not an AI lab

US export controls on Nvidia hardware - including the blocking of H200 imports in January 2026 and the closure of third-country cloud loopholes in May 2026 - were intended to stifle Chinese AI progress. Instead, these policies have acted as a catalyst for domestic independence.

Chinese GPU and AI chipmakers captured 41% of the local accelerator market in 2025, up from near-zero in 2022.

Xiaomi is the clearest example of this loop closing. The simultaneous arrival of the MiMo-V2.6 model and the V900 silicon

  • Xiaomi's own AI chip - represents a company that can now train and serve frontier models on hardware it builds. The $2.62M training cost figure implies they were not using expensive cloud H100 time at rack rates. If the V900 handles the training run at a fraction of the cost of equivalent Nvidia compute, the economics of every future MiMo release look different.

That said, the benchmark table here is vendor-published. The caveats are real: vendor benchmarks, comparison tables that omit the strongest closed models, and genuine gaps on some evaluation categories. Independent reproduction takes time. The Artificial Analysis composite is the most credible external number available right now, and it places MiMo-V2.6-Pro above every other open-weight model. What it does not tell you is how the model performs on your specific task distribution, in your specific agent loop, with your specific context length.

Picking a frontier-tier open-weight model for an agent loop today
Without Beagle
Kimi K3 at 44 on AA Index - $3.00/$9.00 per million tokens, same MIT weights story, smaller omnimodal gap
With Beagle
MiMo-V2.6-Pro at 46 on AA Index - $0.435/$0.87 per million tokens, same MIT license, but needs 8× H200 or an API call to Xiaomi's platform

Xiaomi MiMo-V2.6: common questions

What is Xiaomi MiMo-V2.6-Pro?

MiMo-V2.6-Pro is Xiaomi's open-weight AI model released September 22, 2026 under MIT. It uses a 1.02 trillion parameter mixture-of-experts architecture with 42 billion parameters active per token, supports text, image, video and audio input, and scores 46 on the Artificial Analysis Intelligence Index - the highest score any open-weight model has reached on that leaderboard.

How does MiMo-V2.6-Pro compare to Kimi K3 and Qwen3.8 Max?

MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index; Kimi K3 and GLM-5.3 both score 44. API pricing is significantly lower: Pro costs about one-seventh of Kimi K3 on input tokens. All three are MIT-licensed. The meaningful difference for agent workloads is MiMo-V2.6-Pro's edge on Terminal Bench and CyberGym evaluations, and its omnimodal input support.

Can I self-host MiMo-V2.6-Pro?

Yes, with significant hardware. The official vLLM recipe requires 8× H200 GPUs (680 GB aggregate VRAM) and a nightly vLLM build because stable vLLM ≤ 0.29.0 cannot load the MXFP4 weights. The SGLang production path uses two nodes with 16-way tensor parallelism. For most teams, starting with the Xiaomi API or a third-party host makes more sense until you've validated the model on your workload.

What is MiMo-V2.6-Flash, and how is it different?

Flash is a smaller, cheaper variant in the same release. It requires 208 GB aggregate VRAM (4× H200 or 4× MI355X) and stores weights in 173 GB on disk. Xiaomi's own evals show it scoring 67.9 on DeepSWE v1.1 against Pro's 71.9, and 52.3 on AutomationBench against Pro's 53.1. For most coding agent loops, the gap is meaningful but not decisive - Flash is the sensible starting point if you're self-hosting.

Is MiMo-V2.6 actually MIT licensed?

Yes. Xiaomi publishes the Pro-RL, Flash-RL, and 9B distillation checkpoints on Hugging Face under MIT, which places no restriction on commercial use, modification, or redistribution. The hosted API variants served through Xiaomi's platform are a separate product with separate terms, but the weights themselves carry no commercial restrictions.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle