Safety

Anthropic scales "frontier threats" red teaming—biological risks tested and mitigations identified

Anthropic reports on its ‘‘frontier threats red teaming’’ programme, describing a recent multi‑month project that evaluated biological misuse risks from advanced language models.

Anthropic scales "frontier threats" red teaming—biological risks tested and mitigations identified

Anthropic published a July 26, 2023 post describing its work on "frontier threats red teaming" — adversarial testing focused on high‑impact risks, including biothreats. The company aims to establish a repeatable approach to identify and reduce risks from advanced language models in domains relevant to national security.

Why this matters

Researchers have long warned that advanced language models could acquire capabilities relevant to national security. Anthropic notes that its CEO, Dario Amodei, raised this topic in recent U.S. Senate testimony. The company also supported and joined commitments announced at the White House on July 21, 2023, which included internal and external security testing of AI systems to address major sources of AI risk such as biosecurity and cybersecurity.

How frontier threats red teaming is done

Anthropic emphasizes that red teaming for frontier threats requires significant time and domain expertise. Key elements of their approach:

  • Work directly with domain experts who have decades of experience to define threat models: which information is dangerous, how pieces of information combine to cause harm, and what accuracy and frequency of outputs would be necessary.
  • Follow a structured research plan in which subject‑matter experts and LLM researchers spend substantial time (for example, 100+ hours) interacting with models to probe and understand their capabilities. Experts learn effective interaction patterns and how to attempt to "jailbreak" models.
  • Build new, automated evaluations grounded in expert knowledge and the tooling to run them reproducibly and at scale.
  • Because findings and methods can be sensitive, conduct tests with trusted third parties and strong information security protections.

Findings from red teaming biology

Over six months, Anthropic spent more than 150 hours with leading biosecurity experts to test the model’s ability to generate harmful biological information, such as designing or acquiring biological weapons. Experts used a bespoke, secure interface to the model without the public deployment’s trust and safety enforcement and developed quantitative evaluations of model capabilities.

Main findings:

  • Current frontier models can sometimes produce sophisticated, accurate, expert‑level knowledge in biological domains. That output is not uniformly frequent across all topics, but it occurs in some areas.
  • Evidence suggests model capabilities increase with model scale. Access to external tools could further enhance biological capabilities.
  • Taken together, Anthropic assesses that unmitigated LLMs could accelerate a bad actor’s ability to misuse biology compared with relying solely on internet access, by making it easier to assemble multiple, chained expert‑level pieces of information. Today these effects are small but growing relatively quickly. Anthropic warns these risks may materialize in the near term (for example, within two to three years) rather than only after five or more years.

The company also found that the research process identifies practical mitigations:

  • Straightforward changes in the training process can meaningfully reduce harmful outputs by improving the model’s ability to distinguish harmful from harmless biological uses (Anthropic points to related techniques such as Constitutional AI).
  • Classifier‑based filters make it harder for an adversary to obtain the chained, expert‑level information needed to cause harm. Anthropic has deployed such filters in its public frontier model.

Anthropic has compiled a list of mitigations across the model development and deployment pathway and intends to continue experimenting with them.

Future research priorities

At project close, Anthropic says it has more experiments and evaluations to run than it began with. A high‑priority repeated experiment is measuring how much faster an LLM could enable the production of harmful outcomes compared with a search engine. Such experiments should cover present and future models, including next‑generation, tool‑using, and multimodal systems.

Given the finding that current frontier models provide warning of near‑term risks, Anthropic urges developers of frontier models to perform more analysis and implement stronger mitigations, and to share findings with responsible industry actors and select government agencies. It also recommends preparing for the potential release of models that have not undergone frontier threats red teaming: absent new mitigation approaches, bad actors might extract harmful biological capabilities from smaller, fine‑tuned, or task‑specific models derived from released base weights.

Scaling, collaboration, and responsible disclosure

Anthropic is scaling up its frontier threats red teaming research team to study future capabilities, build scalable evaluations, and develop mitigations. The company is briefing governments and other labs on its findings and piloting a responsible disclosure process so labs and stakeholders can report risks and mitigations to relevant actors.

Anthropic encourages other groups — particularly labs and independent third‑party evaluation organizations — to run similar assessments and offers to support stakeholders interested in doing this work.

Conclusion

Anthropic’s empirical work supports the view that frontier threats red teaming for national‑security‑relevant domains is timely and necessary. Current models show early signals of risks that appear to be growing, creating a window to evaluate and mitigate nascent threats before subsequent model generations with tool access or multimodal capabilities increase those risks. The company calls for cross‑sector collaboration, independent third‑party evaluations, and wider sharing of mitigations with appropriate safeguards.