The UK Artificial Intelligence Safety Institute (AISI) found in its latest tests that OpenAI's publicly available GPT-5.5 performed essentially the same as Anthropic's more restricted Mythos Preview on cybersecurity evaluations. The results were reported by Ars Technica.
Why the tests mattered
Anthropic last month limited access to its Mythos Preview model to “critical industry partners,” citing elevated cybersecurity risks. That decision raised questions about whether Mythos Preview represented a unique breakthrough or whether similar capabilities are present across contemporary large language models.
Test setup: CTFs and complex scenarios
AISI evaluated both models on 95 different Capture the Flag (CTF) tasks. CTFs are cybersecurity competition simulations in which participants solve isolated challenges ranging from web vulnerabilities to cryptography and digital forensics. AISI noted that while CTFs are training-focused and not equivalent to live corporate penetration tests, they model relevant methodologies.
The institute’s evaluation covered tasks such as reverse engineering, exploitation of web vulnerabilities and cryptographic challenges. It also ran bespoke, more complex trials including the 32-step data‑exfiltration simulation called “The Last Ones” (TLO) and a high‑difficulty industrial control scenario dubbed “Cooling Tower.”
Key results and figures
- On the highest, “expert” level challenges GPT-5.5 averaged 71.4% success, while Mythos Preview averaged 68.6% — a difference within the margin of error.
- In a particularly hard task requiring a disassembler for a Rust-based binary, GPT-5.5 solved the challenge in 10 minutes 22 seconds without human intervention, at an API cost of about $1.73.
- In the 32-step TLO data‑exfiltration test, GPT-5.5 succeeded 3 out of 10 trials, while Mythos Preview succeeded 2 out of 10; no single model had previously completed this test successfully.
- On the “Cooling Tower” test, which simulates disrupting an industrial plant's control systems, GPT-5.5 failed just like all prior models; to date no model has completed this scenario.
AISI’s interpretation and implications
AISI concluded that the cybersecurity capabilities demonstrated by Mythos Preview likely do not reflect an isolated, model‑specific breakthrough. Instead, they appear to be byproducts of broader improvements in autonomy, reasoning, and programming ability across modern models. In other words, the risks and capabilities are not necessarily confined to one closed model.
Sam Altman, CEO of OpenAI, described the marketing around limited releases as “fear-based marketing” in a Core Memory podcast interview, while also acknowledging that truly dangerous models may emerge in the future and will need different release strategies.
Industry responses
OpenAI launched a Trusted Access for Cyber pilot in February, allowing security researchers and companies to register to study models for defensive purposes. The company recently restricted access to a cyber‑focused GPT-5.4-Cyber via that list, and Altman has said initial access to GPT-5.5-Cyber will likewise be limited to “critical cyber defense professionals.”
Summary
AISI’s CTF-based testing shows that the public GPT-5.5 model achieves performance close to Anthropic’s restricted Mythos Preview on the evaluated cybersecurity tasks. However, challenges that simulate physical or complex industrial impacts — like the “Cooling Tower” test — remain unsolved by current models. The findings suggest managing future risks will require careful policy on model access and oversight beyond simply locking down a single model.


