Mistral Large 4 scores 93% on Cybench. Claude Opus 5.5 and GPT-6 Astra score near zero on the same test - not because they fail it, but because they refuse to attempt it. That single detail explains more about this launch than any leaderboard position.
Mistral's new flagship, nicknamed Le Chonk, is a natively multimodal mixture-of-experts with about 1 trillion total parameters and 49 billion active per token, in public preview since October 6, 2026.
The preview API is available through Mistral Studio, with weights scheduled for release by the end of October 2026. Until those weights ship, no one outside Mistral can reproduce a single benchmark, inspect the architecture, or read the license terms. That is not a reason to ignore the launch - it is a reason to read the numbers carefully.
What the Mistral Large 4 benchmarks actually show
The launch packs two very different stories into one announcement, and the gap between them is the whole post.
On the Vals.ai broad industry index - which weights finance, coding, legal, and tax tasks together - Vals reports Large 4 at 48.05% ±1.11, ranking 32nd of 44 models in the displayed index. That is a competitive open-weight result, not a frontier-leading one. In the Artificial Analysis Intelligence Index, which aggregates ten benchmarks across different domains, the model scores 38 points but still trails leading closed models like Claude Opus 5.5 by a wide margin.
The narrower story is more interesting. On the Harvey Legal Agent Benchmark, it scores 15.83%, ahead of GPT-6 Astra, Claude Sonnet 5.5, Claude Opus 5.5, and several other competitors in that particular test.
Large 4's separate Legal Research Bench result is 31.73%, ranking 38th of 74, which further cautions against converting one strong agent benchmark into a universal legal superiority claim. One specialist benchmark can sit in the top ten while a broader evaluation places the same model near the bottom third. Both numbers are true.
Mistral's published benchmark table includes DeepSWE v1.1 at 61.7%, SWE-Atlas-QnA at 59.4%, Terminal-Bench 4 at 28.3%, Coding Agent Index at 49.8%, AutomationBench at 59.9%, Cybench at 93%, and Lakera B3 attack resistance at 93.3%. Most of these are Mistral-cited figures from Artificial Analysis or from Mistral's own evaluations. For now, there are no independent benchmarks based on the open weights that could confirm them. The weights are not yet downloadable; independent replication is not yet possible.
The cybersecurity angle - and where it comes from
This is where the launch is genuinely novel and where the framing deserves the most scrutiny.
On the Artificial Analysis Cyber Index, an independent evaluation of how well AI models find and fix security flaws in real software, it ranks among the top five models globally and leads open-weight models developed outside China. On one of the index's tests, which asks a model to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82%, the highest of any model.
Mistral's headline cybersecurity claim is that Claude Opus 5.5 and GPT-6 Astra score near zero on that reproduce-and-patch test because they refuse. That is not a performance failure - it is a safety policy. US closed models increasingly refuse security-research prompts that involve vulnerability reproduction. Mistral's main pitch is cybersecurity work that closed US models refuse to do.
That distinction matters operationally. If you run a penetration testing firm, a threat intelligence team, or a red team function, a model that actually attempts the task is categorically more useful than one that declines it - regardless of where both sit on a general leaderboard. But the same willingness to engage with dual-use prompts is also the reason despite strong performance on cyber benchmarks, Mistral reports that ML4's average refusal rate on malicious cyber prompts from JailbreakBench, StrongREJECT, and AgentHarm is higher than all other open-weight models tested. They are trying to thread both: engage with legitimate security research, reject actual attack assistance. Whether the line holds in production is something only deployment testing will answer.
The VRAM math that most coverage skips
Mistral hardware planning is dominated by GPU VRAM, and the Mixture-of-Experts models add an important nuance: an MoE model must hold all of its experts in VRAM even though only a few activate per token.
Large 4 has 1.05 trillion total parameters, with 49 billion active per forward pass. The active-parameter count drives compute cost and speed. The total-parameter count drives memory. The rule of thumb is roughly 0.5 GB of VRAM per billion parameters under 4-bit quantization; full FP16 doubles that requirement. Applied to 1.05 trillion parameters:
| Quantization | Estimated VRAM | Minimum cluster |
|---|---|---|
| 4-bit (INT4) | ~525 GB | 7× H100 80GB |
| 8-bit (INT8) | ~1,050 GB | 14× H100 80GB |
| FP16 | ~2,100 GB | 27× H100 80GB |
These are estimates; Mistral has not yet published the hardware requirements. Weights are promised for around October 27, 2026; the license and hardware requirements are unpublished. The model is a 1T-parameter MoE far beyond most mid-range clusters, so nothing is plannable or procurable yet. For comparison, the VRAM floor for 4-bit quantization of Mistral Large 3 at 675B parameters is approximately 338 GB, necessitating at least six H100 80GB GPUs. Large 4 is 56% bigger in total parameters, so the cluster requirement scales accordingly.
The practical read: self-hosting Large 4 is a multi-GPU H100-class commitment. Teams without that infrastructure will use the API at $1.36 per million input tokens and $4.18 per million output tokens at list price, though Mistral's model page currently shows half that - $0.68/$2.09 - without saying why or for how long.
What the license gap means right now
License details have not yet been named. Reports suggest a custom Mistral license rather than Apache 2.0, which could include commercial thresholds or use restrictions. Read it before building a product on it.
This is not a minor footnote. The difference between Apache 2.0 and a custom license with revenue gates or MAU caps can be significant. Apache 2.0 models allow full commercial use with no MAU cap and no revenue gate. Mistral Large 3 shipped under Apache 2.0. Qwen3.8-Max's license, by contrast, lets a hosting business run until the licensee and its affiliates clear $50M in aggregate revenue, while the Qwen Community License covering certain Qwen variants requires a separate agreement for any inference resale at all, with no revenue floor. Large 4's terms could land anywhere on that spectrum.
A teammate like Beagle can pull benchmark numbers from a launch post in seconds, but the license terms need a human - and ideally legal counsel - before any product decision.
ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters, and will be available across multiple regions including a European deployment operated independently by Mistral under European law. For EU teams with data residency requirements, that is a meaningful structural advantage over US-hosted closed models - but it is an argument about where inference runs, not about the quality of the outputs. The two questions are separate.
Mistral Large 4 benchmarks: common questions
How does Mistral Large 4 rank against other open-weight models?
On Vals AI's broad industry index it scores 48.05%, ranking 32nd of 44 models. On the narrower Harvey Legal Agent Benchmark it scores 15.83%, ranking 6th of 75. The wide gap between those two ranks reflects a genuine specialist strength rather than a scoring anomaly - legal and cybersecurity post-training appears deliberate.
Can I self-host Mistral Large 4 now?
No. As of October 6, 2026, only hosted API access is available. Mistral says open weights will follow by the end of October 2026, with reports citing dates between October 27 and 31. The license has not been named, and hardware requirements have not been published.
Why does Large 4 beat GPT-6 Astra on cybersecurity benchmarks?
Mistral's headline cybersecurity result is that Claude Opus 5.5 and GPT-6 Astra score near zero on the reproduce-and-patch test because they refuse. The score gap is partly a willingness gap, not purely a capability gap. That matters: if your use case is legitimate security research, the distinction is irrelevant. If you are comparing general coding capability, the cybersecurity benchmark is not the right signal.
Is the Mistral Large 4 API price stable?
The model page currently shows $0.68/$2.09 per million tokens, which is half the list price of $1.36/$4.18, without explanation of why or for how long. Preview pricing on frontier releases frequently changes at the open-weight launch. Treat the current rate as provisional.
When will independent benchmark results exist for Mistral Large 4?
The late-October weights release is when real independent evaluation becomes possible. Preview benchmarks are company-reported; independent evaluation only becomes possible once the weights are downloadable, likely between October 27 and 31, 2026. Until then, the published numbers are sourced from Mistral, Artificial Analysis evaluations that Mistral selected, and third-party evals that Mistral chose to highlight.