In recent months major AI developers have introduced vetted-access programs and strict usage guardrails intended to prevent model misuse by malicious actors. Those protections are increasingly being reported as obstacles by legitimate network defenders and offensive cybersecurity researchers.
Government actions and vendor responses
In June, the U.S. government imposed export controls on Anthropic’s Mythos and Fable models. The move was at least partially prompted by a report that claimed some guardrails could be bypassed, enabling misuse of the models for cyberattacks. Anthropic has publicly positioned Mythos as a particularly sensitive system that should only be available to carefully vetted users.
Some restrictions have since been relaxed: Fable 5 returned to general availability on July 1, while Mythos 5 has been reintroduced only to vetted U.S. organizations as part of an ongoing government review.
Vetted programs and researchers’ criticism
Gatekeeping is not unique to Anthropic. Both Anthropic and OpenAI run programs that let researchers apply for reduced-restriction access — notably OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program. Many security researchers argue these programs and the underlying guardrails hamper their ability to discover and validate previously unknown vulnerabilities.
Mark Dowd, a well-known security researcher, told a cybersecurity podcast: “it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.” Dowd, who has spent decades finding and selling zero-day vulnerabilities to Western governments, acknowledged his perspective may be shaped by his line of work.
When guardrails block legitimate testing
Chris Anley, chief scientist at NCC Group, said asking an AI model to attempt to exploit a bug is often essential to confirm it’s a real vulnerability worth fixing. If a guardrail simply refuses to answer, that can impede defenders as much as attackers. “‘Fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” Anley explained. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.”
When researchers hit such barriers, they sometimes switch to open-source models that have no guardrails.
Paolo Stagno, chief technology officer at CrowdFense — a company that develops and sells unknown vulnerabilities to government agencies — agreed that AI firms often treat customers like “children who need babysitting” with vetted programs and restrictions. Stagno said his team uses frontier models only for reverse engineering and avoids cloud-based AI for vulnerability discovery or exploit construction to prevent leaking sensitive data or having it absorbed into future training runs. Instead they use locally run open-source models for those steps.
Giuseppe Cali, a security researcher who finds zero-days and develops exploits, said guardrails do not hinder him because he does not rely on AI for offensive work. He uses AI for initial reverse engineering and to build supporting tools but retains the actual bug discovery and weaponization himself. “I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” Cali said.
Inconsistent restrictions push researchers to foreign models
An anonymous researcher at a smartphone-component manufacturer said his employer is not part of Anthropic’s Cyber Verification Program, and as a result the company’s tools are of little use for vulnerability research because the guardrails are too strict. Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, added that guardrails can be inconsistent and vary day to day, even within vetted programs. That leads practitioners to spend time ‘negotiating’ with models rather than analyzing vulnerabilities.
Thompson warned this dynamic pushes responsible researchers toward foreign open-source models — for example, Chinese models like GLM that can be downloaded and run locally without vetting. “You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he said. He argued that such a shift may be more harmful than beneficial.
Rather than tightening restrictions, Thompson called on frontier AI labs to broaden responsible access and to enforce accountability for those who abuse tools. Otherwise, he warned, defenders risk falling behind. “There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” Thompson said. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”
What to watch going forward
The core question is how to balance preventing malicious use of AI models with enabling legitimate offensive security research that helps defenders find and fix vulnerabilities. Current vetted programs and strict guardrails reduce abuse risk but can also create operational friction for researchers and encourage migration to less-regulated tools. Observers will be watching whether providers increase program transparency, broaden responsible access, or whether government regulation continues to shape who can use frontier models and under what conditions.



