Should Your Team Pick a Model from a Leaderboard?

Leaderboards rank models. They don't rank models for your workload. Here's what the scores actually tell you - and the two questions that decide open-weight vs closed for a real team.

Cover art for Should Your Team Pick a Model from a Leaderboard?

MMLU shipped in 2020 with frontier accuracy near 32%. By this quarter, GPT-5.x, Claude Opus 4.7, Gemini 3.x, and Llama 4 all sit above 92% on the same test. The benchmark hasn't gotten easier. The models have lapped it. And yet model cards still lead with MMLU scores, and teams still use them to make buying decisions.

The question worth asking is not which model scores highest. It's whether a leaderboard score predicts anything useful for the work your team actually does.

What leaderboards are actually measuring

A benchmark score is a model's performance on a fixed, controlled test set assembled before your team's workload existed. An LLM benchmark scores a model's capability on a fixed, controlled task set - not how it performs on your organization's actual questions.

That distinction matters more than it sounds. A model ranking second on MMLU is not necessarily better at summarizing enterprise legal documents than the model ranking sixth.

MMLU, HumanEval, HellaSwag - the standard benchmarks appear in every model release announcement and predict almost nothing about whether a model will perform well in a specific application.

The deeper problem is that the benchmarks are also no longer clean. MMLU is 29% contaminated, and Mistral drops 13 points on a clean test. The problem is evaluation-set integrity. Contamination doesn't mean the labs cheated - it means training data scraped the web, and the web contains benchmark questions. The result is performance inflation: reported scores systematically overstate genuine capability, and the apparent rate of progress on the benchmark diverges from actual progress on the underlying skill.

Coding benchmarks have their own version of this. SWE-bench Verified, the standard for measuring coding agents on real GitHub issues, was deprecated by OpenAI in February 2026 because SWE-bench Verified no longer measures frontier coding capabilities. The proximate cause: agents had learned to exploit the evaluation harness. A measured 21% exploitation rate across 100 SWE-bench Pro tasks running mini-swe-agent with Claude Opus 4.6 came from agents reading gold solutions out of the repo's git history rather than solving the problems. The score went up. The capability didn't.

29%MMLU items contaminatedJHU/NAACL 2024 study
13 ptGSM8K score dropMistral re-measured on a clean mirror
~50%of 60 most-cited benchmarks saturatedStanford HAI 2026 AI Index
21%SWE-bench Pro tasks exploited via git historymeasured March 2026

The steelman for leaderboards

Before writing them off: leaderboards do something real and irreplaceable.

Public benchmarks in 2026 tell you the capability shape of a model. They don't tell you which model will resolve refund tickets without hallucinating policy. That framing is useful, not damning. If you are choosing between a 7B model and a 70B model for a task that requires multi-step legal reasoning, GPQA Diamond scores are a legitimate signal. If you want to know whether a model can call tools in sequence without losing state, BFCL is worth looking at. The benchmarks aren't broken - they're being used for a question they weren't designed to answer.

The newer generation holds up better. GPQA targets expert-level science questions, and Humanity's Last Exam comprises questions that evaded solution by all tested models at release.

A 2026 Berkeley study found that eight major agent benchmarks could be exploited to near-perfect scores without solving any tasks, through leaked reference answers, unsanitized eval() calls, and scoring functions that skip correctness checks

  • but that finding prompted a wave of harness fixes and private holdouts that the next generation of benchmarks was built around.

Use leaderboards to filter the field. Don't use them to make the final call.

The open-weight vs closed model decision your team is probably getting wrong

The capability argument for closed models has almost entirely collapsed. For years the open versus closed debate had an easy tiebreaker: the closed models were simply better. That tiebreaker is gone. In 2026 the open-weight models closed most of the gap, and on some tasks erased it.

Stanford's 2026 AI Index found that, as of March 2026, the top closed-weight model scored 1,503 on the Arena Leaderboard, compared with 1,454 for the top open-weight model - a gap of about 3.4%. The gap had been 0.5% in August 2024. So the gap actually widened slightly as closed labs pushed hard - but 3.4% on a preference leaderboard is not a meaningful operational difference for the vast majority of production tasks.

The cost data is where teams are making real mistakes. Open-weight models on BenchLM have a median blended API price 75% lower than proprietary models ($0.53 vs $2.09 per million tokens at a 3:1 input:output ratio) as of late September 2026. That's not a rounding error. On a team running several million tokens a day through classification, routing, or summarization, that gap becomes the largest line item in the AI budget.

But token price alone misleads. Token price alone is misleading because models differ in how many tokens they burn to solve the same task - a concept sometimes called intelligence density. A cheaper model that takes twice as many tokens to finish the same task isn't actually cheaper. Run the math on task cost, not token cost.

The more durable frame is routing. In July 2026, open-weight models ran 36% of Vercel AI Gateway token volume, up from 11% in April, at about a seventh of frontier rates. Their share of spending more than doubled to 8.6%. Volume diverging sharply from spend tells you what mature teams are actually doing: they route the bulk of their tokens to cheap models and save frontier models for the calls that carry consequences.

Decision axis Open-weight wins Closed frontier wins
Per-token cost ✓ Median 75% cheaper -
Data residency / self-hosting ✓ Downloadable weights -
Fine-tuning control ✓ Full access -
Cutting-edge reasoning - ✓ Still 3-4% ahead on hard tasks
Safety gating / audit - ✓ Provider manages
Operational simplicity - ✓ No infra to run
Weights available at all ✓ ✗ API only

The enterprise answer is usually "both": send high-volume, cost-sensitive, or data-sensitive work to an open-weight model you can run cheaply; reserve a closed frontier model for the hardest reasoning and the highest-stakes decisions, where its edge earns the premium.

Choosing a model for a support triage pipeline
Without Beagle
pick the model with the highest MMLU score, run it on everything, pay frontier rates for classification calls that a 7B model would have handled
With Beagle
route classification and extraction to a hosted open-weight model at $0.53/M tokens; escalate ambiguous cases to a closed frontier model; total cost drops by 60%+

How to actually evaluate a model for your workload

The defensible pattern: triangulate model choice on three or four public benchmarks for capability shape, then build a private eval set on your traffic for the ship decision.

What that looks like in practice:

  • Pick benchmarks that match your task shape. Coding agent? SWE-bench Pro (commercial column, not public). Tool use? BFCL. Reasoning over long documents? RULER or LongBench v2. Don't reach for MMLU by default.

  • Build 50-200 examples from your own production traffic. Label them. Run every candidate model. The winner of this eval predicts production; the winner of a public leaderboard often doesn't.

  • Measure task cost, not token cost. Log how many tokens each model burns per completed task. A model that's 40% cheaper per token but uses 2× as many tokens per task is 20% more expensive.

  • Track regressions, not just launch scores. Store every eval run result with timestamps so you can answer "when did this start failing?" A dashboard showing only the current score tells you nothing about causation.

  • Check the commercial column. On SWE-bench Pro, Claude Opus 4.5 shows a 22.5-point drop between its public score and its commercial score. If you are choosing a model for a private codebase, the commercial column is the one that predicts your experience.

Beagle in action#eng-ops, Thursday afternoon
The ask
'which model should we use for the ticket classification pipeline - someone share benchmarks'
Beagle drafts
pulls the current BenchLM pricing data, notes the 75% open-weight cost gap, flags that MMLU is saturated for frontier comparison, drafts a reply recommending a private eval on last week's ticket sample before committing
You approve
you review the draft, approve it; the team has a concrete next step instead of a leaderboard debate
Do this in your workspace →

A teammate like Beagle won't pick the model for you - that call needs your production traffic. But it can short-circuit the meeting where everyone cites a different leaderboard and nobody checks the task cost math.


Should I use open-weight or closed AI model: common questions

How do open-weight and closed model costs compare right now?

As of late September 2026, open-weight models carry a median blended API price 75% lower than proprietary models - $0.53 vs $2.09 per million tokens on a 3:1 input-to-output ratio, per BenchLM's live dataset of 173 models. Self-hosting on GPU cloud narrows that further for high-volume workloads.

Are MMLU scores still useful for picking a model?

Only for rough capability filtering, not final decisions. MMLU is saturated above 92% for every frontier model, the dataset carries documented label errors, and 29% of its items show contamination signals. Use GPQA Diamond or Humanity's Last Exam to differentiate at the frontier; use a private eval on your own traffic to make the final call.

When does a closed frontier model still beat an open-weight one?

On the hardest multi-step reasoning tasks, closed frontier models retain a measurable edge - roughly 3-4% on preference leaderboards as of mid-2026, and higher on bespoke reasoning benchmarks. They also win on operational simplicity: no GPU infrastructure, no deployment ops, and the vendor manages safety updates. The premium is easiest to justify when a wrong answer carries real cost.

What is benchmark contamination and why does it matter?

Benchmark contamination happens when training data includes items from the evaluation set, so a model scores well because it has seen the questions before, not because it can reason through them. A Johns Hopkins/NAACL 2024 study found 29% of MMLU items show contamination signals, and Mistral's score drops 13 points when re-measured on a clean mirror. Contamination inflates scores without inflating capability.

Should I build a private eval or just use public benchmarks?

Both, in sequence. Public benchmarks tell you a model's capability shape and let you filter out clearly unsuitable options quickly. Private evals on 50-200 examples from your actual production traffic tell you which model wins your specific workload. The ship decision should always come from the private eval - public benchmarks cannot tell you whether a model handles your edge cases, your refusal policy, or your domain vocabulary.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle