Safety

AI-generated text

Anthropic collaborated with US CAISI and UK AISI to test and harden AI safeguards

Anthropic says it worked with the US Center for AI Standards and Innovation (CAISI) and the UK AI Security Institute (AISI) over the past year to test and strengthen its safeguard systems.

Anthropic collaborated with US CAISI and UK AISI to test and harden AI safeguards

In an announcement dated September 12, 2025, Anthropic described a year-long collaboration with the US Center for AI Standards and Innovation (CAISI) and the UK AI Security Institute (AISI), government bodies created to evaluate and improve AI security. What began as voluntary consultations evolved into an ongoing partnership in which CAISI and AISI teams were granted access to Anthropic’s systems at multiple stages of model development to enable continuous testing.

Why government partners were involved

CAISI and AISI brought capabilities from national-security-related domains — including cybersecurity, intelligence analysis, and threat modeling — combined with machine learning expertise to assess attack vectors and defence mechanisms. According to Anthropic, this external scrutiny helped them harden systems against sophisticated misuse scenarios.

As part of their agreements, both organisations evaluated several iterations of Anthropic’s Constitutional Classifiers — a defence used to detect and prevent jailbreaks — on models such as Claude Opus 4 and Claude Opus 4.1, both before and after deployment, yielding findings that informed stronger safeguards.

Vulnerabilities discovered and responses

Government red teams identified multiple classes of vulnerabilities. Anthropic says its technical teams addressed these issues or changed architectures to mitigate the underlying vulnerability classes. Notable findings included:

  • Prompt injection vulnerabilities: testers used hidden instructions to manipulate models. Specific annotations (for example, falsely asserting that human review had occurred) could bypass classifier detection; Anthropic reports these vulnerabilities were patched.

  • Stress-testing safeguard architectures: red teamers developed a sophisticated universal jailbreak that evaded standard detection methods. Rather than applying a one-off fix, Anthropic restructured its safeguard architecture to address this broader vulnerability class.

  • Cipher-based attacks: encoded requests using ciphers, character substitutions, and other obfuscation techniques were used to evade classifiers. These discoveries prompted improvements enabling detection of disguised harmful content regardless of encoding.

  • Input and output obfuscation attacks: universal jailbreaks were found that fragmented harmful strings into seemingly benign components within broader contexts tailored to Anthropic’s specific defenses. Identifying these blind spots led to targeted filtering improvements.

  • Automated attack refinement: testers built automated systems that iteratively optimized attack strategies, turning less effective jailbreaks into effective universal jailbreaks; Anthropic used these iterations to further strengthen safeguards.

Evaluation approach and risk methodology

Beyond specific vulnerabilities, CAISI and AISI contributed to Anthropic’s broader security approach. Their external perspective on evidence requirements, deployment monitoring, and rapid response helped pressure-test assumptions and identify areas where additional evidence or adjustments to threat models were needed.

Practical lessons from the collaboration

Anthropic highlights several lessons about working effectively with government research and standards bodies to improve model safety:

  • Comprehensive model access increases red-teaming effectiveness: providing pre-deployment safeguard prototypes and multiple system configurations — from unprotected to fully safeguarded models — allowed testers to develop and refine attacks progressively.

  • Extensive documentation and internal resources: Anthropic shared details of safeguard architectures, documented vulnerabilities, safeguard reports, and granular content policy information, which supported more targeted testing.

  • Real-time safeguards data accelerates discovery: giving red teams access to classifier scores enabled faster refinement of attack methods and more focused exploratory research.

  • Iterative testing reveals complex vulnerabilities: sustained collaboration, including daily communication channels and frequent technical deep-dives during critical phases, enabled external teams to develop deep expertise and uncover more complex weaknesses.

  • Complementary, multi-layered security works best: public bug bounty programs provide volume and diversity of reports, while specialised government expert teams find complex, subtle vectors; together they create a more robust defensive ecosystem.

Conclusion

Anthropic argues that making powerful AI models secure and useful requires both technical innovation and new forms of industry–government collaboration. The company notes that other AI developers are also engaging with CAISI and AISI and encourages more firms to work with such bodies and share lessons. Anthropic also thanked the technical teams at US CAISI and UK AISI for their rigorous testing and collaboration, saying this work materially improved the security of their systems and advanced methods for measuring the effectiveness of AI safeguards.