Read the AISI Open-Weight Safety Benchmark Before You Trust the Gap

The UK AI Security Institute just published its first public measurement of the open-weight vs. closed-model frontier gap on cyber tasks. Here's what the numbers actually say - and why the shrinking gap matters more than it looks.

Cover art for Read the AISI Open-Weight Safety Benchmark Before You Trust the Gap

The UK's AI Security Institute published a number this week that deserves a closer read than it's been getting: open-weight models now trail the closed-model cyber frontier by four to seven months, down from six to ten through most of 2025. That sounds like a reassuring lag. It isn't.

AISI evaluated GLM-5.2 and DeepSeek V4-Pro and found they perform similarly to frontier closed models released four to seven months before them - a narrower gap than the six to ten months measured through most of 2025. The evaluation is the first time AISI has published this comparison publicly, and the methodology is worth understanding before you use the number for anything.

What AISI actually measured - and what it didn't

On a set of 70 evals for specific, narrow cyber capabilities, GLM-5.2 is closest to Claude Opus 4.6, which was released 4.3 months earlier, while DeepSeek V4-Pro sits somewhere between Claude Opus 4.5 and GPT-5, released in November and August 2025 respectively. Those are narrow tasks: discrete skills like vulnerability research, reverse engineering, and cryptography, graded across four difficulty tiers.

The second method is harder to run and harder to spin. AISI's "The Last Ones" cyber range simulates a 32-step attack on a corporate network with four subnets and about 20 hosts, and AISI estimates a human expert would need roughly 20 hours to complete it. On that test, the gap is wider. GLM-5.2 matches Opus 4.5 on The Last Ones - released roughly seven months before it - while DeepSeek V4-Pro falls below Sonnet 4.5 from the same period.

That distinction matters for how you interpret "four to seven months." The four-month figure is for discrete skills. The seven-month figure is for end-to-end, multi-step autonomous attack chains - the scenario that is actually hard to defend against.

AISI's own tracking shows closed-model cyber performance was doubling every 4.7 months as of February 2026, itself accelerated from an eight-month doubling time measured in November 2025. A lag of four months measured against a frontier that doubles every five months is a different problem than a four-month lag against a static target.

Why the safeguard finding is the real story

AISI found that their open-weight model evaluations were largely unimpeded by safeguards. Of the two models tested, DeepSeek V4-Pro occasionally refused narrow cyber tasks, but this was easily circumvented by a small number of repeat attempts at refused tasks.

This is what separates open-weight risk from closed-model risk in a way that the month-lag framing obscures. Critics see risk in open models because anyone can download, modify, and run them without oversight - and once released, users can remove safety guardrails, share copies freely, and run them on private systems beyond anyone's control. The gap in time is the gap before those capabilities are available without any gatekeeping. When it was ten months, that felt like a meaningful buffer. At four to five months on discrete tasks, it is less comfortable.

AISI frames this plainly: these findings indicate a narrow window before today's frontier cyber capabilities may become widely accessible without safeguards.

4-7 monthsopen/closed frontier gap nowdown from 6-10 months in 2025
70narrow cyber tasks testedacross 4 difficulty tiers
32 stepsin The Last Ones cyber range~20 human-expert hours
4.7 monthsfrontier capability doubling timeas of February 2026

The cost asymmetry nobody is pricing in

There's a second dimension to the open-weight story that the AISI report surfaces indirectly: at open-weight prices, frontier-adjacent capability is cheap.

At advertised first-party prices, a fixed 100-million-token run costs about $85 with Anthropic's hosted Claude Opus 4.6 - and an estimated $46 with GLM-5.2.

DeepSeek V4-Pro's standing API price makes it cheaper still. Specifically: DeepSeek V4-Pro costs $0.435 per million input tokens and $0.87 per million output tokens.

That cost gap looks even wider if you self-host. The weights are public. Users can host open-weight models privately with no data flowing back to providers, customize them, cut costs, and rely on a foundation that providers can't change or shut down. The point is not that self-hosting is easy (it isn't - see our post on what the hardware actually costs). The point is that the economics of frontier-adjacent capability are now radically different depending on whether you're buying API access or downloading weights.

Model API input ($/1M) API output ($/1M) AISI gap vs closed frontier
GLM-5.2 $1.40 $4.40 ~4 months (narrow tasks), ~7 months (ranges)
DeepSeek V4-Pro $0.435 $0.87 ~5 months (narrow tasks)
Claude Opus 4.6 (closed) ~$3.00 ~$15.00 frontier

At first-party API rates, DeepSeek V4-Pro costs $0.435/$0.87 per million input/output tokens compared to GLM-5.2's $1.40/$4.40 - making DeepSeek V4-Pro roughly 3.2× cheaper on input and 5× cheaper on output. Both are substantially cheaper than the closed frontier they're chasing.

What Kimi K3 will tell us next

Moonshot AI released their latest flagship model Kimi K3 on July 16th - a 2.8 trillion parameter MoE model with weights scheduled to release July 27th. AISI has said directly that it intends to test Kimi K3 on the same basis once its weights are released.

The key fact is that either the open-to-closed or American-to-Chinese model performance gap has been reduced from the debated six to nine months to something shorter, say three to five months. If K3's weights test where its API benchmarks suggest, the four-month figure in the current AISI report may not hold for long.

The honest read on the current report: the result was found despite rapid advancements AISI observed in frontier cyber capabilities up to February 2026, but it is not predictive of whether future open-weight models will replicate the more recent jumps delivered by Mythos Preview and GPT-5.5. The AISI evaluation is a snapshot, not a trend line.

Beagle in action#security-review, 10:40am
The ask
'anyone have a read on what the AISI open-weight report actually means for our model policy?'
Beagle drafts
pulls the AISI primary source and the two model pricing pages, drafts a one-paragraph summary with the key gap numbers and the safeguard caveat
You approve
you approve the draft; the context lands in the thread without someone spending 45 minutes hunting primary sources
Do this in your workspace

For teams using open-weight models in production - whether for coding agents, document analysis, or anything that touches sensitive data - this report is a useful forcing function. The question isn't whether your use case involves cyber tasks. It's whether your model-selection policy accounts for the fact that safeguards on open-weight releases are currently at the level of "one or two retries."

AISI open-weight safety benchmark: common questions

What did the AISI open-weight benchmark actually find?

AISI found that recent open-weight models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released four to seven months before them - a narrower gap than the six to ten months measured through most of 2025. The four-month figure applies to discrete cyber tasks; the seven-month figure applies to multi-step autonomous attack simulations.

Is the open-weight frontier gap closing or widening?

It is closing in calendar months, but the frontier itself is moving. AISI's tracking shows closed-model cyber performance was doubling every 4.7 months as of February 2026 , which means a four-month lag measured against an accelerating frontier translates to more absolute capability than the same lag did a year ago.

Why do open-weight model safeguards matter differently than closed-model safeguards?

Open models whose weights anyone can download can have their safety guardrails removed, be shared freely, and run on private systems beyond anyone's control. With closed APIs, the provider retains enforcement leverage. With open weights, the release is permanent and unrestricted.

Which models did AISI test and why?

AISI tested GLM-5.2 and DeepSeek V4-Pro, selected because they were candidates to lead open-weight cyber capability at the time of release.

AISI intends to test Kimi K3 on the same basis once its weights are publicly released.

Should teams stop using open-weight models based on this report?

AISI is explicit that open-weight models also offer real benefits: users can host them privately with no data flowing back to providers, customize them, cut costs, and rely on a foundation that providers can't change or shut down. The report argues these competing concerns need to be balanced - not that open weights are categorically unsafe.

Keep reading