Epoch AI ran the numbers in May 2026: the best open-weight model on any given day matched the capability of the best closed model from roughly four months earlier. That is a tighter gap than most people assume, and a slightly wider one than last year. Both facts matter if you are picking a model for a production workflow.
Since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months in Epoch's Capabilities Index. The average ECI gap was 8 points - similar to the gap between GPT-5 and GPT-5.5. For context, that gap is slightly larger than the one Epoch identified in October 2025, which put the average lag at three months between January 2023 and October 2025.
The four-month number is precise and sourced. It is also only half the picture.
What the 4-month figure actually measures
Epoch calculates the average time it takes for the best open-weight models to catch up with state-of-the-art performance according to their Epoch Capability Index - a composite measure that captures performance across many benchmarks. The methodology matters: they are measuring a time-equivalent, not a raw score delta. A four-month lag means the best open model today performs like the best closed model did in June 2026. It does not mean the open model is 4/12ths as capable.
There is a subtlety buried in that framing that most coverage skips. Across 17 selected benchmarks, private evaluations show a gap of 8-10 months, almost twice the gap on public ones at 4-6 months. That divergence is the non-obvious part. The benchmarks open models train toward - MMLU, HumanEval, coding tasks - get saturated faster because model developers can see them. Private evaluations, which no training team can goodhart against, show a gap nearly twice as large.
What that means practically: if you are selecting a model based on public leaderboard scores, you are systematically underestimating how much capability is left on the table. A closed model that scores similarly to an open one on Chatbot Arena may still outperform it on your actual internal task, especially if that task involves long-horizon planning, tool use across many steps, or domain-specific reasoning that doesn't appear in public evals.
The cost picture flips when you include hosted open-weight APIs
Much of the open-weight appeal is economic. Open models are more cost-effective: GLM-5.2, Qwen3.5, and DeepSeek-V4 perform better than frontier models at the same evaluation cost. That is real. But the comparison gets murkier once you separate three different ways to run an open-weight model.
The model weights are free to download, but running them in production is not. You still pay for inference - per-token on a hosted API, GPU hours plus operations if you self-host, or hardware and setup for local deployment. "Open weights" lowers licensing and lock-in costs; it does not eliminate the cost of serving predictions.
Against a budget API like DeepSeek V4 Flash at around $0.14 per 1M input tokens, the breakeven for a single H100 sits at roughly 5.7 billion tokens per month. Against a premium GPT-5-class API at about $5 per 1M, the crossover can arrive near 256 million tokens per month - but that figure assumes 60-70% sustained GPU utilization.
Most teams do not sustain 60-70% GPU utilization. Effective cost-per-token scales inversely with load. At roughly 10% utilization, your real cost per token can be about 10× the headline GPU rate - enough to make an idle H100 more expensive per token than a premium frontier API.
The cleaner path for most teams: a hosted open-weight API is the lowest-friction way to start. Providers such as Together AI, Groq, and Fireworks AI run the GPUs, keep the models updated, and charge per million tokens. You get the open-weight price advantage without owning the utilization problem.
Where open-weight models genuinely win today
The capability gap is real. It is not a reason to default to closed models for every workload.
When you look at the cost-performance Pareto frontier, open models dominate the low- and mid-cost portion of the frontier. For teams running high-volume, structured tasks - classification, extraction, summarization, JSON parsing - an open-weight model served through a hosted API is often the right call on both cost and quality. The 4-month capability lag matters far less when the task is "extract these ten fields from a document" than when it is "design a multi-step research plan and execute it."
Since January 2025, the capability gap between leading open-weight and closed models has narrowed. That trend is real too. Alibaba's Qwen models gained traction, occupying the top spot for an open-weight model on Chatbot Arena as of August 2025. And in August 2025, OpenAI released its first open-weight models since GPT-2 in 2019: gpt-oss-120b and gpt-oss-20b. The release of gpt-oss changed the calculus in a specific way: suddenly the "open-weight" bucket included a model from the lab that also produces the closed frontier. That compression matters.
Where closed models still lead outright:
Long-horizon agentic tasks. Long-horizon workflows involving tool use, planning, and repeated decision-making expose larger differences between models, especially on complex tool environments and multi-step orchestration.
Multimodal tasks. Leading proprietary frontier models often perform strongly on the hardest reasoning, multimodal, and long-horizon agentic evaluations, although rankings vary by benchmark and task.
Narrow but high-stakes domains. Claude Sonnet 4.6 was the sole closed-source model on the Pareto frontier in one oncology decision-making evaluation - achieving a 14.5% score advantage at 5.5× the cost. When the accuracy ceiling matters more than the cost floor, closed models still hold it.
A teammate like Beagle, which runs inside Slack and Teams, operates on short-horizon tasks - answering questions, drafting replies, surfacing context - where the open/closed gap shrinks considerably. The question of which model brain to route those requests through matters, but it is far less decisive than for a coding agent working on a 40-step plan.
How to actually use the 4-month figure
Treat it as a task classifier, not a verdict. The question is not "is the open-weight model four months behind?" but "does my specific task sit in the part of the capability curve where four months is decisive?"
Three questions that sharpen the answer:
- What does failure look like? A wrong summary is recoverable. A wrong multi-step agent action may not be. Higher stakes push toward closed.
- What is the volume? At millions of calls per day, even a small price-per-token difference compounds fast. Open-weight (hosted API) wins on economics for high-volume, lower-stakes tasks.
- Is there a private-eval equivalent for your task? Benchmark scores do not show consistency, retry rates, or performance on a company's real workloads. Independent evaluators compare models across coding, reasoning, tool use, and agentic tasks - but teams should still test representative workloads because performance varies by model, task design, latency requirements, and cost constraints.
The honest answer to "open-weight or closed?" in late 2026 is: open-weight, hosted, for structured and moderate-complexity tasks; closed for long-horizon agentic work where the private-eval gap is most likely to bite you. And run your own eval on a representative sample before you commit either way.
Open-weight vs. closed model: common questions
How far behind are open-weight models compared to closed models?
Since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months in Epoch's Capabilities Index. The average ECI gap was 8 points. The gap is wider - 8-10 months - on private evaluations that haven't been trained against.
Is self-hosting an open-weight model cheaper than a closed API?
It depends on volume and utilization. Against a budget API like DeepSeek V4 Flash at ~$0.14 per 1M tokens, the self-hosting break-even for a single H100 is roughly 5.7 billion tokens per month. Against a GPT-5-class API at ~$5 per 1M, the crossover arrives near 256 million tokens per month - at 60-70% sustained utilization. Most teams never hit sustained utilization that makes self-hosting win.
Do open-weight models beat closed models on cost-per-result?
On many mid-complexity tasks, yes. Open models like GLM-5.2, Qwen3.5, and DeepSeek-V4 perform better than frontier closed models at the same evaluation cost. The exception is the top of the accuracy range, where closed models still define the frontier.
What tasks show the biggest open vs. closed model gap?
Long-horizon workflows involving tool use, planning, and repeated decision-making expose the largest differences. Short, structured tasks - extraction, classification, summarization - show smaller gaps and are where open-weight models offer the best value.
Should I pick a model based on Chatbot Arena or public leaderboards?
Leaderboard scores are useful signals but systematically understate the gap for real workloads. Private evaluations show a gap of 8-10 months - almost twice the 4-6 month gap measured on public benchmarks. Run a sample of your actual tasks before committing.