On February 23, 2026, OpenAI's Frontier Evals team stopped reporting SWE-bench Verified scores. Not because their model was losing. Because the benchmark had stopped meaning anything - the 500-issue coding test that had become the unofficial scoreboard for "can the machine code yet" was, in their assessment, too contaminated to carry signal. That decision is a useful place to start when thinking about open-weight versus closed models, because most teams are still making this choice using numbers that share the same problem.
The honest state of the landscape: recent open models like GLM-5.2 and DeepSeek V4-Pro perform similarly to closed frontier models released four to seven months before them
- a gap measured by the UK's AI Safety Institute on its own internal cyber capability evals, not a vendor slide. That trails the closed frontier by a narrower margin than the six to ten months measured through most of 2025. The gap is real but shrinking. The decision your team actually faces is not "which is smarter." The open-weight models closed most of the gap, and on some tasks erased it - so the enterprise decision is no longer about which is smarter. It is about how much control you need over cost, data, and change. Most teams are still arguing about capability when they should be deciding about control.
What the benchmarks actually measure (and what they don't)
A leaderboard number is a measurement of a model under specific conditions, not a general capability score. LLM benchmark methodology in 2026 has a credibility problem: the headline number on a leaderboard is frequently the least reliable thing about it. Every widely cited static benchmark is contaminated to some degree, and identical model weights can score 10-20 percentage points apart depending on the evaluation harness.
The SWE-bench Verified leaderboard listed 100 models on June 16, 2026 - but only one result was independently verified; the other 99 were submitted by the vendors themselves. When you read that a model scored 82% on SWE-bench, you are largely reading a self-report. Independent benchmarking analysis notes that any Verified score above 80% "warrants scrutiny about harness and tool access" - and that when the leading dozen models cluster between 80% and 95% on a self-reported benchmark, the rank order stops carrying much signal.
Frontier models gained 30 percentage points in a single year on Humanity's Last Exam, a benchmark built to be hard for AI. Evaluations intended to be challenging for years are saturated in months, compressing the window in which benchmarks remain useful for tracking progress.
The contamination problem is structural, not accidental. Data contamination is a serious issue: the performance of LLMs on competitive programming and SWE-Bench tasks has been shown to degrade over time, indicating the possibility of older problems being contaminated due to public exposure on the internet. The successor benchmark, SWE-bench Pro, tries to address this: it has 1,865 total tasks spanning 731 public, 858 held-out, and 276 commercial problems, across 41 repositories in Python, Go, TypeScript, and JavaScript, with tasks averaging 107.4 lines changed across 4.1 files.
The cost arithmetic teams get wrong
Cost is the loudest argument for open-weight models, and also the most misread one. The sticker price of an API call is not the cost of running inference.
Off-peak, DeepSeek V4.1 Flash is $0.15 per million input tokens (cache miss) and $0.60 per million output tokens. V4-Pro is $0.66 input / $1.98 output per million tokens.
Even at those rates, V4-Flash output runs roughly 23 to 45 times cheaper than GPT-5.5 at equivalent context lengths. That spread is real and worth taking seriously for high-volume workloads.
But self-hosting open weights does not make inference free. The crossover between using DeepSeek's hosted API and running your own cluster lands around 831 million tokens a day on spot GPU pricing for a budget INT4 quantized V4-Flash deployment. Most teams are not near that threshold. The infrastructure cost - provisioning, patching, scaling, monitoring - is invisible in the per-token comparison and very visible in your on-call rotation.
| Decision axis | Open-weight advantage | Closed advantage |
|---|---|---|
| Data residency | Full control; data stays in your network | Depends on vendor DPA and region |
| Cost at volume | Dramatically cheaper above ~100M tokens/day | No infrastructure overhead at low volume |
| Time to production | Weeks of infra work | Ship the day you get an API key |
| Fine-tuning | Fine-tune or distill the weights you own | Prompt engineering only (mostly) |
| Model stability | Pinned to a specific checkpoint | Vendor can update silently |
| Frontier reasoning | 4-7 months behind leading closed models | Still ahead on hardest long-horizon tasks |
The actual decision axes
The practical question is not open versus closed. It is which workload properties force the choice.
Data residency is the clearest forcing function. If a single prompt leaving your network is unacceptable - patient records, financial data, classified material, or a strict data-residency mandate - self-hosted open-weight is usually the only option that fully satisfies the requirement.
EU data residency rules are driving a measurable move toward self-hosting, and the EU AI Act became fully applicable on 2 August 2026.
Self-hosting an open-weight model resolves the residency question cleanly: data does not leave your infrastructure, there is no third-party processing agreement to manage, and audit trails are entirely within your control.
Volume and workload shape matter more than capability rank. If your agent is doing classification, extraction, summarization, routing, or light generation at scale, open-weight models can handle this with minimal quality difference.
Use a premium frontier model for architecture, ambiguous debugging, and high-risk changes. Use a cheaper open-weight hosted model for draft patches, test generation, and refactors. Use a local model for private code exploration, offline work, and low-risk repetitive tasks.
The operational burden is not optional. Self-hosted open weights remove the vendor API boundary, so teams need their own logging, access control, cost tracking, and audit trails. A team without dedicated ML platform engineering that self-hosts a frontier-scale open-weight model - a 304B parameter model like DeepSeek V4-Flash requires serious GPU infrastructure - is taking on work that the closed API price already pays for elsewhere.
Model stability is underrated. Closed APIs update silently. When you pin to a specific DeepSeek or GLM checkpoint, that checkpoint does not change. For agentic pipelines where prompt behavior needs to be predictable across thousands of runs, pinning to a downloaded weight is a meaningful guarantee that no API contract provides. A teammate like Beagle, routing tasks across models in Slack and Teams, benefits from that predictability at the step level - some steps need a pinned, deterministic open-weight model; others can tolerate a managed API.
Open weight vs closed AI models: common questions
What is the capability gap between open-weight and closed models right now?
The UK AI Safety Institute measured it at 4-7 months as of July 2026: the best open-weight models perform comparably to closed frontier models released four to seven months earlier. That gap has narrowed from 6-10 months in 2025. On some specific tasks - long-context coding, structured output - the gap is effectively zero.
Can I trust AI benchmark leaderboard scores when comparing models?
Treat them as directional filters, not measurements. Most scores on widely cited leaderboards like SWE-bench Verified are vendor self-reported. The evaluation harness alone can shift a score by 10-20 percentage points for the same model weights. Use leaderboards to generate a shortlist, then run your own eval on your actual data and tasks before committing.
When does self-hosting an open-weight model make financial sense?
The crossover against a cheap hosted API like DeepSeek V4.1 Flash sits around 831 million tokens per day on a budget GPU cluster. Below that, the API is almost always cheaper once you account for infrastructure, engineering time, and operational overhead. The financial case for self-hosting is strongest when compliance forces the decision first, not cost.
What open-weight models are production-grade right now?
As of September 2026, the production-credible open-weight options are DeepSeek V4.1 Flash (MIT-licensed, 552B parameters, 8B active, strong cost-performance), GLM-5.3-Flash (320B total, 18B active, multimodal), and Qwen3-Coder (480B-A35B, scores 73.4 on SWE-bench Verified while fitting on a single MacBook for the smaller variants). Choose by workload shape and license, not headline benchmark rank.
Does open-weight mean open source?
Not always. Open-weight means the trained parameters are publicly downloadable. The training data, architecture details, and code may or may not be published. License terms vary significantly: DeepSeek V4 Flash uses MIT, Qwen3-Coder uses Apache 2.0, but some larger Qwen variants require revenue-sharing agreements above $50M annual revenue. Read the license before you ship a derivative.