On June 12, 2024, Anthropic published a detailed account of the red teaming approaches it has used to probe its AI systems and offered recommendations for developing shared practices. The post is intended to inform other companies conducting red teaming, policymakers interested in practical testing, and organizations that aim to test AI technologies.
What is red teaming?
Red teaming is an adversarial testing practice meant to reveal vulnerabilities in a technological system. Researchers and developers use a variety of red teaming techniques today, each with distinct trade‑offs. Anthropic stresses that the current absence of standardized red teaming practices complicates objective comparisons of different AI systems’ safety.
The company argues that establishing norms now is important so organizations can manage present risks and prepare for future threats as model capabilities expand.
Overview of red teaming methods covered
- Domain‑specific, expert red teaming
- Trust & Safety: Policy Vulnerability Testing (PVT)
- National security: frontier threats red teaming
- Region‑specific: multilingual and multicultural red teaming
- Using language models to red team
- Automated red teaming
- Red teaming in new modalities
- Multimodal red teaming
- Open‑ended, general red teaming
- Crowdsourced red teaming for general harms
- Community‑based red teaming for general risks and system limitations
The following sections examine each method’s advantages and challenges.
Domain‑specific, expert red teaming
This approach engages subject matter experts to identify and evaluate risks within their fields, bringing deeper, context‑specific insight to potential vulnerabilities.
Policy Vulnerability Testing (PVT) for Trust & Safety
High‑risk threats that could cause severe harm or societal impact require sophisticated red teaming and collaboration with external experts. Anthropic uses a form of qualitative, in‑depth testing called Policy Vulnerability Testing (PVT) together with outside specialists across policy areas covered by its Usage Policy. The post names partners such as Thorn (on child safety), Institute for Strategic Dialogue (on election integrity), and Global Project Against Hate and Extremism (on radicalization).
Frontier threats red teaming for national security risks
Anthropic has continued developing evaluation techniques for so‑called frontier threats—areas that could present consequential national security risks. Their frontier work focuses on Chemical, Biological, Radiological, and Nuclear (CBRN) risks, cybersecurity, and autonomous AI risks. External experts help both test systems and co‑design new evaluation methods. Depending on the threat model, external red teamers may work with standard deployed versions of Claude in “real‑world” settings or with non‑commercial versions that use different mitigations.
Multilingual and multicultural red teaming
Most red teaming starts from English and a U.S. perspective. To address this representational gap, Anthropic conducts red teaming in other languages and cultural contexts. Public sector capacity‑building can encourage local communities to test models on language and topics relevant to them. Anthropic cites a project with Singapore’s Infocomm Media Development Authority (IMDA) and AI Verify Foundation that covered four languages (English, Tamil, Mandarin, and Malay) and Singapore‑relevant topics.
Using language models to red team
Language models can be employed to generate adversarial inputs automatically, augmenting manual testing and enabling broader and faster coverage of potential attack vectors.
Automated red teaming and the red team/blue team dynamic
As models improve, Anthropic explores using models themselves to carry out automated red teaming. They describe a red team/blue team loop: a model (red team) generates attacks likely to elicit target behaviors, and another model (blue team) is fine‑tuned on those red‑teamed outputs to become more robust against similar attacks. Repeating this cycle can surface new attack vectors and strengthen model defenses.
Red teaming in new modalities and multimodal systems
Testing systems that accept inputs beyond text—such as images or audio—can uncover new failure modes. The Claude 3 family is multimodal: while it does not generate images, it can accept visual inputs (photos, sketches, charts) and produce text responses. These capabilities introduce risks (fraud, child safety threats, violent extremism, etc.), so Anthropic’s Trust & Safety team performed pre‑deployment red teaming for image and text risks and engaged external red teamers to assess refusal behaviors and mitigations.
Open‑ended and community‑based red teaming
Anthropic began red teaming research in mid‑2022 with crowdworkers in a controlled research setting. Public events and competitions—such as DEF CON’s AI Village and the Generative Red Teaming (GRT) Challenge—have broadened participation. Anthropic noted the enthusiasm and creativity of participants from diverse backgrounds and hopes such efforts will attract a wider set of people to AI safety work.
From qualitative red teaming to quantitative evaluations
Anthropic outlines an iterative pathway from ad hoc, qualitative red teaming toward automated, quantitative evaluation. The process starts with subject matter experts defining a threat model and probing the model to elicit harmful behavior. As red teamers refine and standardize inputs, language models can be used to generate hundreds or thousands of variants to increase coverage quickly. Anthropic has applied this iterative approach in frontier threats work and in Policy Vulnerability Testing for election integrity and plans to extend it to other threat models.
Policy recommendations
To support adoption and standardization of red teaming, Anthropic recommends that policymakers:
- Fund organizations like the National Institute of Standards and Technology (NIST) to develop technical standards and common practices for safe and effective AI red teaming.
- Finance the creation and operation of independent government and nonprofit bodies that can partner with developers to red team systems across domains; for national security risks, much expertise will reside in government agencies.
- Encourage the growth of a market for professional AI red teaming services and establish a certification process for organizations conducting red teaming according to shared technical standards.
- Urge AI companies to allow and facilitate third‑party red teaming by vetted (and eventually certified) outside groups, and develop standards for transparency and model access to enable this under safe conditions.
- Encourage AI firms to tie red teaming practices to clear policies that set conditions for continued scaling or release of new models (for example, commitments similar to a Responsible Scaling Policy).
Conclusion
Anthropic frames red teaming as a vital method for detecting and mitigating AI risks. The post catalogs multiple techniques applicable to different threat models and use cases and calls for collaboration to iterate on these methods and work toward common standards. According to Anthropic, investing in red teaming is one important element among several needed to develop AI systems responsibly and with robust safeguards.



