Safety

Anthropic finds rapid gains in cyber and bio capabilities but not yet at national-security thresholds

Anthropic reports that its Frontier Red Team observed rapid improvements in frontier AI capabilities across cybersecurity and biology during 2024–2025, with Claude models reaching undergraduate-level performance in many Capture The Flag tasks and exceeding expert baselines on some virology evaluations.

Anthropic finds rapid gains in cyber and bio capabilities but not yet at national-security thresholds

Anthropic published findings from its Frontier Red Team work covering the past year (post dated March 19, 2025). The post summarizes what the company has learned about the trajectory of potential national-security risks from frontier AI models. Anthropic assesses that models are showing “early warning” signs of rapid progress in key dual-use capabilities—particularly in cybersecurity and certain areas of biology—often approaching or exceeding undergraduate-level performance. However, the company states current models remain below thresholds at which they would consider national-security risks to be substantially elevated.

High-level context

Anthropic emphasizes that real-world risk depends on factors beyond raw AI capability: physical constraints, specialized equipment, human expertise, and practical implementation challenges continue to be important barriers. With those caveats, the company describes observed capability trends across focal domains.

Cybersecurity

Anthropic characterizes 2024 as a “zero to one” year for cyber capabilities. In Capture The Flag (CTF) exercises—controlled challenges that require finding and exploiting software vulnerabilities—Claude progressed from roughly high-school level to undergraduate level in about a year. On a set of publicly available Intercode CTFs, performance improved from solving less than a quarter of challenges to nearly all in under a year.

The latest model, Claude 3.7 Sonnet, continued that progress. On Cybench—a public benchmark that uses CTF problems to evaluate large language models—Claude 3.7 Sonnet solves about a third of challenges within five attempts, up from about five percent with Anthropic’s frontier model at the same point last year.

Improvements appear across multiple categories of CTF tasks (pwn, web, and crypto). Nevertheless, the models still trail expert human operators on some skills: they struggle with reverse engineering binary executables and with reconnaissance and exploitation in a network environment unless given assistance.

Working with outside experts at Carnegie Mellon University, Anthropic ran experiments on more realistic cyber ranges of roughly 50 hosts that test multi-stage operations requiring reconnaissance and lateral movement. In those network settings, models are not yet capable of autonomous success. However, when equipped with a toolkit developed by researchers (called Incalmo), Claude (and other LLMs) could, with simple instructions, replicate an attack similar to a known large-scale theft of personally identifiable information from a credit reporting agency.

Anthropic says this evaluation infrastructure helps provide warning if autonomous capabilities improve, and can also inform AI-assisted cyber defense work.

Biosecurity

Anthropic continued its biosecurity evaluations. The company reports rapid progress in the models’ biological understanding: within a year Claude moved from underperforming world-class virology experts on a troubleshooting evaluation to comfortably exceeding that baseline (the VCT evaluation designed by SecureBio).

Biological capabilities remain uneven. Internal tests show models approaching human-expert baselines on tasks relevant to wet-lab work—understanding protocols and manipulating DNA and protein sequences—and the latest model exceeds expert baselines on cloning workflows. The models are still worse than human experts at interpreting scientific figures.

To understand how these improving but uneven skills translate into biosecurity risk, Anthropic ran small, controlled studies of weaponization-related tasks and consulted world-class biodefense experts. In one experimental study, the most recent model provided some uplift to novices compared to participants without model access; however, even the highest-scoring plan from a participant using a model contained critical errors that would cause real-world failure.

External red-teaming judgments were mixed: some experts saw improved knowledge in specific weaponization aspects, while others judged the number of critical planning failures too high for an end-to-end attack to succeed. Overall, Anthropic concludes its models cannot reliably guide a novice malicious actor through key practical steps of bioweapon acquisition and use at present. Given the rapid improvement, the company continues to invest heavily in monitoring biosecurity risks and in mitigations such as its recent work on constitutional classifiers.

Strategic partnerships and external testing

Anthropic notes that pre-planned evaluation and capability thresholds let it move quickly while managing risk. The company’s models have undergone pre-deployment testing by the US AI Safety Institute and the UK AI Security Institute (AISI); those assessments informed Anthropic’s understanding of Claude 3.7 Sonnet’s national-security-relevant capabilities and influenced the model’s AI Safety Level (ASL) determination.

Anthropic also led a first-of-its-kind partnership with the National Nuclear Security Administration (NNSA), part of the US Department of Energy, in which the NNSA is evaluating Claude in a classified environment for nuclear and radiological knowledge. Because of the sensitivity of nuclear-weapons-related information, red-teaming in this project was carried out directly by the government. Anthropic shared insights from CBRN risk identification and mitigation approaches that NNSA adapted for the nuclear domain.

Looking ahead

Anthropic highlights the importance of internal safeguards (for example, its Responsible Scaling Policy), independent evaluation bodies (AI Safety/Security Institutes), and appropriately targeted external oversight. The company plans to scale up to more frequent tests with automated evaluation, elicitation, analysis, and reporting.

Anthropic warns that as models gain better extended thinking and planning abilities, toolkits like Incalmo may become less necessary and models may perform cybersecurity tasks better out of the box. Based in part on the biological research described, Anthropic believes its models are approaching capability thresholds that would trigger AI Safety Level 3 safeguards, and it is investing to ensure those measures are ready in time. The company calls for deeper collaboration between frontier AI labs and governments to improve evaluations and risk mitigation across these domains.

If readers are interested in contributing directly, Anthropic states it is hiring.