Safety

AI-generated text

Anthropic and U.S. DOE/NNSA co-develop AI classifier to detect nuclear misuse

Anthropic has partnered with the U.S.

Anthropic and U.S. DOE/NNSA co-develop AI classifier to detect nuclear misuse

Anthropic announced that it has co-developed an automated classifier with the U.S. Department of Energy (DOE)’s National Nuclear Security Administration (NNSA) to identify whether nuclear-related conversations are concerning or benign. The effort moves beyond risk assessment toward an operational monitoring capability for nuclear-proliferation risks associated with advanced AI models.

Background of the partnership

Anthropic notes that nuclear technology is inherently dual-use: the same physics that enables reactors can also be misapplied for weapons. As AI models become more capable, the company says it is important to monitor whether those models could provide users with dangerous technical knowledge that threatens national security.

Anthropic began working with the DOE/NNSA in April (2025) to assess its models for nuclear proliferation risk. The new work is a step further: instead of only evaluating risks, the partners have built a tool designed to detect potential misuse in real time.

What was built and how it performs

Together with DOE national laboratories, Anthropic co-developed an AI classifier that, in preliminary testing, distinguished concerning from benign nuclear-related conversations with 96% accuracy. Anthropic reports that the classifier has already been deployed on traffic for its Claude model as part of a broader system for identifying misuse.

Early deployment data and next steps

Initial deployment data suggest the classifier performs well on real Claude conversations. Anthropic plans to share its approach with the Frontier Model Forum, an industry body for frontier AI companies, with the goal of providing a blueprint other AI developers could use to implement similar safeguards in partnership with the NNSA.

Why this matters

Anthropic frames the project as a first-of-its-kind example of public–private collaboration to address national-security risks from frontier AI models. The company says such partnerships combine complementary strengths of industry and government to make AI systems more reliable and trustworthy.

Timing and source

The announcement was published on August 21, 2025. Anthropic’s Frontier Red Team blog post contains further details about the NNSA partnership and the safeguards development.