Safety

UK AI Security Institute: gap between open-weight and closed models' cyber abilities is narrowing

The UK’s AI Security Institute (AISI) analyzed differences in cybersecurity capabilities between leading proprietary models and open-weight models and found the performance gap has decreased in 2026.

UK AI Security Institute: gap between open-weight and closed models' cyber abilities is narrowing

The UK government’s AI Security Institute (AISI) has published its first public analysis comparing cybersecurity capabilities of leading open-weight models with those of proprietary, closed-weight frontier models. AISI reports that in 2026 the performance gap has narrowed: recent open models are closer to the closed frontier than they were through most of 2025.

What AISI measured and the results

AISI evaluated models on a suite of 70 narrow, specific cyber capability tasks. On these tests GLM-5.2 is the nearest open model to Claude Opus 4.6, which was released 4.3 months earlier. DeepSeek V4-Pro ranks between Claude Opus 4.5 and GPT-5; GPT-5 was released in August 2025 and Claude Opus 4.5 in November 2025.

Overall, AISI finds that GLM-5.2 and DeepSeek V4-Pro perform similarly to closed frontier models that were released roughly 4–7 months earlier. That is a smaller lag than the 6–10 month gap AISI measured across most of 2025.

Longer-horizon cyberrange tasks show a larger gap

The gap grows for long-horizon cyberrange exercises where models must chain multiple capabilities to complete a full hacking operation. On a cyberrange called The Last Ones, AISI reports GLM-5.2 reaches as far as Opus 4.5 (a model released less than 7 months earlier), while DeepSeek V4-Pro falls below Sonnet 4.5 (a sub-frontier model released 7 months earlier). AISI notes the gap on these chained tasks is larger than on the narrow evaluations.

AISI’s commentary suggests this pattern aligns with industry observations that open-weight models can match proprietary models on surface-level or narrow tasks, but sometimes lack the broader generalization that distinguishes frontier closed models — a phenomenon informally referred to as the industry term "big model smell."

Why this matters

AISI emphasizes the operational implication: as the distance between the controllable proprietary frontier and the openly diffused frontier narrows, defenders have less time to prepare before frontier-level cyber capabilities may be widely accessible without the safeguards employed by proprietary companies. In AISI’s words, this implies a short window for cyber defenders to adapt before such capabilities become available outside the same controls.

Next steps

AISI says it intends to test the Kimi K3 model on the same basis once its weights are publicly released.

Summary

The UK AI Security Institute’s measurements show that 2026 open-weight releases GLM-5.2 and DeepSeek V4-Pro have closed much of the performance gap with closed frontier models on narrow cyber tasks, though larger differences remain on long-horizon, chained exercises. The institute warns that the shrinking gap shortens the time defenders have to respond before powerful capabilities spread without the protections used by proprietary providers.