
The UK’s AI Security Institute, or AISI, says leading open-weight models have narrowed the cyber-capability gap with top proprietary AI systems to roughly four to seven months, a shorter lag than earlier in 2025 and a notable warning for security teams counting on time to adapt. In AISI’s assessment, models such as GLM-5.2 and DeepSeek V4-Pro can now match the performance levels that some closed systems reached only months earlier, while costing much less to run.
That combination matters beyond benchmarking. According to AISI’s analysis, lower-cost open models with downloadable weights reduce the practical barriers to scaling offensive cyber workflows, while also making provider-level safety controls harder to enforce. For builders and enterprise buyers, the report sharpens a tension that has defined the open-model debate: the same attributes that make open-weight systems attractive for private deployment and customization can also make misuse harder to contain.
AISI said this is its first public assessment of how far top open-weight models trail leading closed models on cyber-related tasks. The institute used two evaluation methods.
The first, called Narrow Cyber Tasks, covers 70 tasks across four difficulty levels. According to The Decoder’s report on the findings, those tasks span vulnerability research, reverse engineering, web exploitation, and cryptography. On that benchmark, AISI found that GLM-5.2, released in June 2026, performed at about the level of Opus 4.6 from February 2026, implying a gap of around four months. DeepSeek V4-Pro, meanwhile, was said to perform around the level of Opus 4.5, which AISI dates to November 2025.
The second method, Cyber Ranges, tests more autonomous behavior in simulated networks rather than isolated tasks. One scenario, “The Last Ones,” models a multi-step attack on a corporate network with four subnets and about 20 hosts. AISI estimated that a human expert would need around 20 hours to complete that exercise.
Results there were less favorable for open models, but still pointed to rapid catch-up. AISI said GLM-5.2 performed about as well as Opus 4.5 in the simulated environment, while DeepSeek V4-Pro landed below Sonnet 4.5. The institute described the Cyber Ranges evidence as weaker than the Narrow Cyber Tasks results because it relies on fewer scenarios and may reflect limits in long-horizon planning rather than pure cyber knowledge.
That distinction is important. For security buyers, the practical risk is not just whether a model can solve a reverse-engineering prompt, but whether it can sustain a complicated sequence of actions across a noisy environment. AISI appears to be saying both capabilities matter, and the gap between open and closed models may depend on which kind of work is being measured.
AISI’s reported cost figures are arguably the most operationally significant part of the release. According to the institute, a 100-million-token Cyber Ranges test cost about $85 using Opus 4.5 or Opus 4.6, about $46 using GLM-5.2, and only $1.19 using DeepSeek V4-Pro.
It also reported per-task cost estimates for reliably solved tasks. Opus 4.6 came in at roughly $15 per task, GLM-5.2 around $6, Opus 4.5 about $12.50, and DeepSeek V4-Pro about $0.28.
Those figures are not the same as real-world attack costs, and AISI did not claim they were. But they do suggest that “good enough” cyber performance may be becoming far cheaper to access, especially when paired with open deployment options. For defenders, that means the threat is not limited to whether open models exactly match the current frontier. If older or slightly weaker capabilities can be run at very low cost and high volume, they may still change attacker economics.
This is one reason the report is relevant to enterprise AI, not only security research. Many organizations evaluating open-weight models focus on privacy, sovereignty, and lower inference cost. AISI’s findings suggest they also need to think about dual-use exposure: a model that is attractive for internal coding, automation, or red-team support may have safety properties that differ sharply from API-based closed alternatives.
AISI’s starkest conclusion is about controls, not raw performance. It said the safety measures on tested open models were largely ineffective. In one example cited by The Decoder, DeepSeek V4-Pro sometimes refused reverse-engineering requests, but retrying was enough to bypass the restriction.
That weakness is not presented as a flaw unique to a single model. AISI’s broader point is structural: safeguards such as monitoring, classifiers, throttling, and user restrictions depend on controlling access to the system. Once weights are released, that control weakens or disappears. Copies can be run privately, modified, or shared beyond the reach of the original developer.
AISI reportedly describes that as a “persistent and irreversible risk of misuse.” That is stronger language than a routine model card warning, and it helps explain why the institute frames the shrinking gap as a policy and defense-timing problem rather than just a leaderboard update.
At the same time, the report does not dismiss open models outright. AISI also recognizes the practical benefits of open-weight deployment: private hosting, lower vendor dependence, customization, and reduced risk that a provider changes access terms or turns a service off. For many companies, those are real advantages. The policy challenge is that the same decentralization that makes open models attractive to buyers also reduces centralized control over harmful use.
The core findings in this story come from AISI’s evaluation as reported by The Decoder; the second source, Tech Times, appears to reflect the same underlying assessment without adding material detail. That means the main evidence base here is a specialist media account of an institutional analysis, not a full public paper reproduced in the source set.
Several caveats matter. First, AISI said the Cyber Ranges results are weaker evidence because they come from fewer test cases. Second, the institute also said the evaluations may slightly underestimate open models at their best because those systems were not specially tuned for the tests. Third, the simulated network environments do not include active defenders or the full messiness of real organizations.
Those limitations cut in different directions. Open-model advocates could argue the reported gap overstates the lead held by closed systems if better prompting or tuning would improve results. Security practitioners, on the other hand, could argue the simulations understate real-world friction for attackers because production networks are defended, monitored, and inconsistent.
There is also a benchmark-design issue beneath the headlines. A model that scores well on Narrow Cyber Tasks may still fail in a genuine multi-stage intrusion because it loses track of the plan, mishandles tools, or makes compounding errors. Conversely, a model that struggles in a long autonomous range may still be very useful to a human operator who breaks work into smaller steps. AISI’s own framing appears to acknowledge that distinction.
For AI builders, the report points to a changing baseline. If open-weight systems like GLM-5.2 and DeepSeek V4-Pro can approach recent closed-model cyber performance at sharply lower cost, product teams building coding assistant, security automation, or AI agents workflows will have more deployment options — but also more responsibility for how those systems are constrained, observed, and segmented.
For enterprises, the practical takeaway is not “avoid open models” so much as “separate cost decisions from risk decisions.” Running an open model in a controlled internal environment can still make sense, especially where data residency and private inference matter. But security teams may need stronger local controls, narrower tool access, more detailed audit logging, and clearer internal-use policies than they would rely on with a hosted API.
For the cyber market, AISI’s timeline framing may be the biggest strategic signal. The institute sees the gap between closed and open models as a temporary preparation window in which defenders using stronger systems can adapt before equivalent capabilities become broadly downloadable. If that window is now four to seven months instead of six to ten, roadmaps for detection engineering, red teaming, and abuse monitoring may need to move faster.
The report also lands amid rapid movement at the high end of the model market. AISI said closed models including Mythos Preview and GPT-5.5 delivered some of the largest gains in AI cyber capability since it began testing. The UK’s National Cyber Security Centre has separately warned that the cyber threat environment is changing quickly. Even if open models do not instantly replicate every frontier jump, the pace of diffusion appears to be accelerating.
One immediate signal is AISI’s planned testing of Kimi-K3, which The Decoder says is expected to have open weights released in late July. AISI reportedly believes current coding benchmarks suggest Kimi-K3 could come closer to today’s frontier models, though potentially at higher cost than other open models.
Beyond that single release, three things matter. First, whether future Cyber Ranges evaluations show open models improving on long-horizon autonomy, not just narrow task competence. Second, whether enterprises adopt stronger local governance patterns for open-weight deployment rather than assuming provider-era safety habits still apply. Third, whether closed-model vendors can preserve a meaningful lead in cyber performance, or whether frontier gains continue to spread into downloadable systems within months.
The notable part of this AISI finding is not simply that open models are improving. It is that the lag appears to be compressing while inference economics improve even faster. In practice, that means the market may not need exact frontier parity for open models to become strategically important in cyber workflows. “Close enough, cheap enough, and uncontrollable enough” is already a serious threshold.
For product teams and buyers, the lesson is to treat open-weight adoption as an infrastructure and governance decision, not just a model-quality decision. The same factors driving interest in enterprise AI autonomy — private hosting, lower cost, and freedom from vendor lock-in — also weaken inherited safety assumptions. As models like GLM-5.2, DeepSeek V4-Pro, and possibly Kimi-K3 close the gap, the real differentiator may shift from raw capability to who can deploy them with reliable controls around tools, data, and abuse.
AISI says open-weight models now trail top closed AI cyber systems by just 4-7 months, shrinking defenders’ lead as costs fall sharply.