Frontier AI models are evolving faster than the established methods used to test and benchmark their cyber capabilities. As a result, policymakers and corporate security teams lack clear, reliable ways to predict what these systems can do and whether they can be deployed safely.
What governments and industry are doing
- Federal agencies have an August 1 deadline to create a classified benchmarking process to assess frontier AI capabilities. The Financial Times reports these standards could appear as soon as this week.
- Anthropic said that with the return of Fable 5 it is developing a standardized benchmark together with Amazon, Google, Microsoft and other partners. Anthropic’s stated approach emphasizes the outcomes and impact of a jailbreak rather than only whether a jailbreak is possible.
Industry initiatives
Even before government action accelerated, industry groups were already rethinking how to measure cyber-capable AI:
- Irregular, a testing lab that works closely with frontier AI labs including OpenAI and Anthropic as well as governments, released a new cyber benchmark in late June. That benchmark measures whether models can perform offensive cyber tasks such as remote code execution, privilege escalation, and reaching restricted networks.
- Other companies and groups, including Wiz and Vals AI, are also developing benchmarks that evaluate how well models carry out similar offensive cyber operations.
- Stanford’s 2026 AI Index warned that “evaluations intended to be challenging for years are saturated in months.”
How testing has changed
Earlier benchmarks typically focused on isolated tasks: predictable, staged hacking challenges or finding previously patched vulnerabilities that were not part of a model’s training data. But agentic and advanced reasoning capabilities in models like Mythos Preview and GPT‑5.5 are rapidly outpacing those tests, complicating efforts to understand what these systems can actually accomplish in realistic environments.
David Slater, co‑founder of AI red‑teaming company Armadin, told Axios: “We're testing maybe the most bare bones fundamentals of capabilities. We are very far away from measuring whether this thing can, in a real environment, do something dangerous.”
Practical experience from red‑teaming
Slater said his company's AI agents surpassed every public cyber benchmark within four weeks by combining additional training and human expertise. By the last quarter of 2025, Armadin concluded public cybersecurity benchmarks were “totally saturated” and “useless.”
He argues that next‑generation benchmarks must assess whether models can carry out longer, more sophisticated cyberattacks and how much effort or cost is required. This includes testing in environments that resemble real production systems to better indicate how quickly models can bypass security controls or move laterally through a network.
Sandbox escape attempts and testing limits
Slater also noted models are improving at attempting to escape sandboxed environments, making it harder for defenders to evaluate them in isolated settings that don't interact with production systems. “The jailbreak attempts are nuts,” he said. “We see this thing trying to escape and get out onto the cloud container that it's running on, using keys that it has access to, to do crazy stuff.”
What to watch next
Attention is focused on how Washington will decide to evaluate the cyber capabilities of U.S. frontier AI models, as leading AI labs push back against the current, largely ad hoc testing process. The design of new benchmarks will be critical to providing a realistic picture of the risks and capabilities of these systems.



