Read Mistral Large 4's Benchmarks Before You Plan Around Them

Mistral Large 4 launched October 6 with striking cybersecurity scores-but ranks 32nd of 44 on the broad Vals index. Here's what the numbers actually mean for teams evaluating it.

Cover art for Read Mistral Large 4's Benchmarks Before You Plan Around Them

Mistral Large 4 scores 93% on Cybench. Claude Opus 5.5 and GPT-6 Astra score near zero on the same test - not because they fail it, but because they refuse to attempt it. That single detail explains more about this launch than any leaderboard position.

Mistral's new flagship, nicknamed Le Chonk, is a natively multimodal mixture-of-experts with about 1 trillion total parameters and 49 billion active per token, in public preview since October 6, 2026.

The preview API is available through Mistral Studio, with weights scheduled for release by the end of October 2026. Until those weights ship, no one outside Mistral can reproduce a single benchmark, inspect the architecture, or read the license terms. That is not a reason to ignore the launch - it is a reason to read the numbers carefully.

What the Mistral Large 4 benchmarks actually show

The launch packs two very different stories into one announcement, and the gap between them is the whole post.

On the Vals.ai broad industry index - which weights finance, coding, legal, and tax tasks together - Vals reports Large 4 at 48.05% ±1.11, ranking 32nd of 44 models in the displayed index. That is a competitive open-weight result, not a frontier-leading one. In the Artificial Analysis Intelligence Index, which aggregates ten benchmarks across different domains, the model scores 38 points but still trails leading closed models like Claude Opus 5.5 by a wide margin.

The narrower story is more interesting. On the Harvey Legal Agent Benchmark, it scores 15.83%, ahead of GPT-6 Astra, Claude Sonnet 5.5, Claude Opus 5.5, and several other competitors in that particular test.

Large 4's separate Legal Research Bench result is 31.73%, ranking 38th of 74, which further cautions against converting one strong agent benchmark into a universal legal superiority claim. One specialist benchmark can sit in the top ten while a broader evaluation places the same model near the bottom third. Both numbers are true.

Mistral's published benchmark table includes DeepSWE v1.1 at 61.7%, SWE-Atlas-QnA at 59.4%, Terminal-Bench 4 at 28.3%, Coding Agent Index at 49.8%, AutomationBench at 59.9%, Cybench at 93%, and Lakera B3 attack resistance at 93.3%. Most of these are Mistral-cited figures from Artificial Analysis or from Mistral's own evaluations. For now, there are no independent benchmarks based on the open weights that could confirm them. The weights are not yet downloadable; independent replication is not yet possible.

The cybersecurity angle - and where it comes from

This is where the launch is genuinely novel and where the framing deserves the most scrutiny.

On the Artificial Analysis Cyber Index, an independent evaluation of how well AI models find and fix security flaws in real software, it ranks among the top five models globally and leads open-weight models developed outside China. On one of the index's tests, which asks a model to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82%, the highest of any model.

Mistral's headline cybersecurity claim is that Claude Opus 5.5 and GPT-6 Astra score near zero on that reproduce-and-patch test because they refuse. That is not a performance failure - it is a safety policy. US closed models increasingly refuse security-research prompts that involve vulnerability reproduction. Mistral's main pitch is cybersecurity work that closed US models refuse to do.

That distinction matters operationally. If you run a penetration testing firm, a threat intelligence team, or a red team function, a model that actually attempts the task is categorically more useful than one that declines it - regardless of where both sit on a general leaderboard. But the same willingness to engage with dual-use prompts is also the reason despite strong performance on cyber benchmarks, Mistral reports that ML4's average refusal rate on malicious cyber prompts from JailbreakBench, StrongREJECT, and AgentHarm is higher than all other open-weight models tested. They are trying to thread both: engage with legitimate security research, reject actual attack assistance. Whether the line holds in production is something only deployment testing will answer.

Beagle in action#security-ops, 10:42am
The ask
'does Large 4 actually handle our vuln-repro workflow or will it just refuse?'
Beagle drafts
pulls the Cybench score, the reproduce-and-patch result, and Mistral's refusal-rate claim from the launch post; drafts a one-paragraph summary with source links
You approve
you approve; the team has the actual numbers in-thread in 30 seconds, not a 10-minute Google session
Do this in your workspace →

The VRAM math that most coverage skips

Mistral hardware planning is dominated by GPU VRAM, and the Mixture-of-Experts models add an important nuance: an MoE model must hold all of its experts in VRAM even though only a few activate per token.

Large 4 has 1.05 trillion total parameters, with 49 billion active per forward pass. The active-parameter count drives compute cost and speed. The total-parameter count drives memory. The rule of thumb is roughly 0.5 GB of VRAM per billion parameters under 4-bit quantization; full FP16 doubles that requirement. Applied to 1.05 trillion parameters:

Quantization Estimated VRAM Minimum cluster
4-bit (INT4) ~525 GB 7× H100 80GB
8-bit (INT8) ~1,050 GB 14× H100 80GB
FP16 ~2,100 GB 27× H100 80GB

These are estimates; Mistral has not yet published the hardware requirements. Weights are promised for around October 27, 2026; the license and hardware requirements are unpublished. The model is a 1T-parameter MoE far beyond most mid-range clusters, so nothing is plannable or procurable yet. For comparison, the VRAM floor for 4-bit quantization of Mistral Large 3 at 675B parameters is approximately 338 GB, necessitating at least six H100 80GB GPUs. Large 4 is 56% bigger in total parameters, so the cluster requirement scales accordingly.

The practical read: self-hosting Large 4 is a multi-GPU H100-class commitment. Teams without that infrastructure will use the API at $1.36 per million input tokens and $4.18 per million output tokens at list price, though Mistral's model page currently shows half that - $0.68/$2.09 - without saying why or for how long.

1.05Ttotal parametersall must sit in VRAM for MoE serving
49Bactive per tokenwhat drives compute speed, not memory
32nd of 44Vals broad index rankvs. 6th of 75 on Harvey Legal Agent

What the license gap means right now

License details have not yet been named. Reports suggest a custom Mistral license rather than Apache 2.0, which could include commercial thresholds or use restrictions. Read it before building a product on it.

This is not a minor footnote. The difference between Apache 2.0 and a custom license with revenue gates or MAU caps can be significant. Apache 2.0 models allow full commercial use with no MAU cap and no revenue gate. Mistral Large 3 shipped under Apache 2.0. Qwen3.8-Max's license, by contrast, lets a hosting business run until the licensee and its affiliates clear $50M in aggregate revenue, while the Qwen Community License covering certain Qwen variants requires a separate agreement for any inference resale at all, with no revenue floor. Large 4's terms could land anywhere on that spectrum.

A teammate like Beagle can pull benchmark numbers from a launch post in seconds, but the license terms need a human - and ideally legal counsel - before any product decision.

Evaluating Large 4 for a security team
Without Beagle
reading Mistral's announcement, noting the 93% Cybench score, assuming it replaces the current workflow
With Beagle
checking that the score reflects task completion not just refusal-avoidance, confirming the license permits your use case, and running your actual vuln-repro prompts against the preview API before weights ship

ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters, and will be available across multiple regions including a European deployment operated independently by Mistral under European law. For EU teams with data residency requirements, that is a meaningful structural advantage over US-hosted closed models - but it is an argument about where inference runs, not about the quality of the outputs. The two questions are separate.

Mistral Large 4 benchmarks: common questions

How does Mistral Large 4 rank against other open-weight models?

On Vals AI's broad industry index it scores 48.05%, ranking 32nd of 44 models. On the narrower Harvey Legal Agent Benchmark it scores 15.83%, ranking 6th of 75. The wide gap between those two ranks reflects a genuine specialist strength rather than a scoring anomaly - legal and cybersecurity post-training appears deliberate.

Can I self-host Mistral Large 4 now?

No. As of October 6, 2026, only hosted API access is available. Mistral says open weights will follow by the end of October 2026, with reports citing dates between October 27 and 31. The license has not been named, and hardware requirements have not been published.

Why does Large 4 beat GPT-6 Astra on cybersecurity benchmarks?

Mistral's headline cybersecurity result is that Claude Opus 5.5 and GPT-6 Astra score near zero on the reproduce-and-patch test because they refuse. The score gap is partly a willingness gap, not purely a capability gap. That matters: if your use case is legitimate security research, the distinction is irrelevant. If you are comparing general coding capability, the cybersecurity benchmark is not the right signal.

Is the Mistral Large 4 API price stable?

The model page currently shows $0.68/$2.09 per million tokens, which is half the list price of $1.36/$4.18, without explanation of why or for how long. Preview pricing on frontier releases frequently changes at the open-weight launch. Treat the current rate as provisional.

When will independent benchmark results exist for Mistral Large 4?

The late-October weights release is when real independent evaluation becomes possible. Preview benchmarks are company-reported; independent evaluation only becomes possible once the weights are downloadable, likely between October 27 and 31, 2026. Until then, the published numbers are sourced from Mistral, Artificial Analysis evaluations that Mistral selected, and third-party evals that Mistral chose to highlight.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle