Three days ago, Xiaomi released MiMo-V2.6-Pro and it immediately topped every open-weight leaderboard that matters. By scoring 46 on the Artificial Analysis Intelligence Index v4.3, it tied Grok 4.7, establishing a new ceiling for open-weight models. That is a real achievement. It is also a number that will mislead your team if you read it wrong.
Here is the thing about that 46: closed models like Claude Opus 5 at maximum reasoning scored 51, and GPT-5.6 Sol scored 47. A five-point gap on a composite index sounds small. Whether it matters for your workload depends entirely on what those five points are measuring - and most teams never check.
Why the classic benchmarks stopped being useful
The benchmarks you have heard of - MMLU, HumanEval, GSM8K - are no longer discriminating. The benchmark landscape has shifted fundamentally. Classic benchmarks like MMLU, HumanEval, and GSM8K are saturated - frontier models score above 90% and no longer differentiate. When the top and the fourth-best model are separated by two percentage points, you are measuring noise from training run variance, not capability.
The benchmarks that shaped the public's mental model of AI progress - MMLU, HumanEval, even GPQA Diamond - have either saturated or contaminated their way out of usefulness at the frontier. The honest reading of 2026 leaderboards starts from a posture of informed scepticism: every static number is suspect, every agentic number is harness-dependent, and every arena rank lives inside a confidence interval that usually overlaps its neighbours.
Contamination makes this worse. OpenAI's own Frontier Evals team reported that models can reproduce original gold patches or problem statements verbatim from the evaluation using nothing but the task ID - an effect that touches all tested frontier models. When the answer is memorisable from an identifier, the score stops measuring software-engineering ability and starts measuring recall.
The Hugging Face Open LLM Leaderboard, once a go-to reference, is gone. It was officially retired in June 2025. Its replacement landscape is fragmented: composite indices, agentic harnesses, arena votes, each measuring a different slice of reality.
What agentic evals actually measure - and where MiMo-V2.6 sits
Agentic evals are now the frontier metric, and they tell a more interesting story. For assessing state-of-the-art models in 2026, look at three things: composite indexes for the overview, agentic and long-horizon evals for differentiation, and autonomy measurements for the trend.
MiMo-V2.6-Pro's performance on these is good but uneven. On DeepSWE v1.1, which measures extended software development work, MiMo-V2.6-Pro scored 71.9 versus 74.0 for Claude Opus 5 and 73.0 for GPT-5.6 Sol. MiMo also trails closed models on Terminal Bench 4.0. Those are the hard, multi-step coding tasks where the closed-model moat still holds.
Where it pulls ahead: on AutomationBench, which covers business task operation, Pro scored 53.1 and Flash 52.3, both surpassing the comparison models. For teams running document-heavy, workflow-style agents - form processing, ticket triage, data extraction - that gap in its favour is the one that actually matters.
The SWE-bench number that any vendor shows you should also be read carefully. SWE-bench scores vary 25 percentage points depending on scaffolding. The same model, with a different agent harness around it, can look dramatically better or worse. You are not just picking a model; you are picking a model-plus-harness combination.
| Benchmark | What it tests | Saturated? | MiMo-V2.6-Pro |
|---|---|---|---|
| MMLU | General knowledge | Yes (>90% frontier) | Not discriminating |
| HumanEval | Single-function coding | Yes | Not discriminating |
| SWE-bench Verified | Multi-file code patches | No, but harness-dependent | 71.9 (vs 74.0 Claude Opus 5) |
| AutomationBench | Business task operation | No | 53.1 - leads closed models |
| AA Intelligence Index | Composite weighted | No | 46 (first open-weight) |
The open-weight gap: four months, not four years
Since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months in the Epoch Capabilities Index, with an average ECI gap of 8 points - similar to the gap between GPT-5 and GPT-5.5. That is a concrete, measurable lag, not a vague hand-wave about open models being behind.
Four months of capability lag is small enough that for most team workloads - support, internal Q&A, data processing, code review - open-weight is already sufficient. The gap is large enough that on hard agentic coding and long-horizon planning, closed models still win.
MiMo-V2.6-Pro sits at the top of the open-weight Artificial Analysis Intelligence Index at roughly $0.13 per task, while the closed frontier scores in the low-to-mid 50s at $3 to $8. That is a 23x to 60x cost difference for a 10% capability gap on composite evals. For high-volume agent loops where you are calling a model thousands of times a day, that arithmetic matters more than the leaderboard rank.
MiMo-V2.6 was trained with a single large-scale reinforcement learning run that Xiaomi livestreamed, reportedly costing around $3.5 million. That training cost is now below what many enterprises spend on a year of closed-model API bills for a single product team. The economics of open-weight are shifting faster than most teams have updated their assumptions.
How to build a quick eval before you commit
The right answer to any model launch is the same: run your own traces before you move production workloads.
Sample 100-200 real inputs from your actual workload. Not curated, not cleaned - whatever your agent or pipeline actually receives.
Pick two or three benchmarks that map to your task type. For a support agent: AutomationBench, multi-turn arena scores. For a coding agent: SWE-bench Verified, Terminal Bench. Running MMLU for a code assistant tells you almost nothing useful.
Run the same harness on the old and new model. The scaffolding has to be identical or the comparison is meaningless.
Measure cost alongside quality. A model ranked #1 on a leaderboard may be many times more expensive per token than the model at #4. Read benchmark results alongside a cost-per-token comparison and a reasoning-effort versus quality breakdown, since extended-thinking modes inflate scores while inflating cost and latency in lockstep.
Check the license. MiMo-V2.6 ships MIT. Its weights can be downloaded for local or hosted deployment, but open weight does not automatically mean open source: training data and training code may remain private, and commercial restrictions can still apply.
Open-weight model benchmarks: common questions
What does the Artificial Analysis Intelligence Index actually measure?
The Artificial Analysis Intelligence Index is a weighted composite across ten evaluation categories: agents, coding, scientific reasoning, and general knowledge, with agents weighted at 34%. It is currently the most cited independent reference because it uses its own methodology rather than relying on vendor submissions. It does not replace workload-specific evaluation.
Is MiMo-V2.6-Pro the best open-weight model right now?
As of September 23, 2026, it scores first among open-weight models on the Artificial Analysis Intelligence Index v4.3. That is the honest framing. It trails closed frontier models on extended software development tasks, and its agentic eval coverage on some third-party harnesses is still sparse. It is the best value in open-weight, not the outright strongest model.
Why do SWE-bench scores vary so much between reports?
HumanEval suffers from training data contamination. SWE-bench scores vary 25 percentage points depending on scaffolding. The agent harness - how context is packaged, how tools are called, how retries are handled - changes the score dramatically. Always check what harness a vendor used before comparing numbers across providers.
Should my team use open-weight or closed models for agents?
For business task automation, document processing, and support workflows, open-weight models at the current frontier (MiMo-V2.6-Pro, Qwen3.8 Max, Kimi K3) are cost-competitive and capable. For hard multi-file coding, long-horizon planning, and tasks where 3-5% accuracy at the margin matters, closed frontier models still hold an edge. The four-month lag in capability is real; decide whether it matters for your specific task before choosing on cost alone.
How do I know if a benchmark is contaminated?
Look for whether the evaluation set is public and static. If it is, assume contamination is possible - models train on large web crawls that include benchmark discussions and solutions. Dynamic benchmarks like LiveBench, which refresh questions, and private held-out sets like those used by Scale AI are more resistant. Vendor-reported scores on public benchmarks are the least trustworthy; independently run scores on private sets are the most.