Safety

AI-generated text

Anthropic publishes a broad framework to assess and manage AI harms

Anthropic on April 21, 2025 published an expanded framework for identifying and mitigating a wide range of potential harms from its AI systems, complementing its Responsible Scaling Policy focused on catastrophic risks.

Anthropic publishes a broad framework to assess and manage AI harms

On April 21, 2025, Anthropic published an exposition of its evolving approach to identifying and mitigating a broad range of harms that could arise from its AI systems. The company positions this framework as a complement to its Responsible Scaling Policy (RSP), which focuses specifically on catastrophic risks, and says a wider perspective is needed to capture the full spectrum of potential impacts.

Why this framework matters

Anthropic argues that rapid advances in AI capabilities call for more systematic, structured ways to consider and manage potential consequences. The framework is intended to help teams communicate clearly, make evidence-based decisions, and develop targeted solutions for both known and emerging harms. The company cautions that the approach is still developing and invites collaboration across the AI ecosystem.

Structure of the framework

The framework assesses potential AI impacts across multiple baseline dimensions and evaluates each dimension using factors such as likelihood, scale, affected populations, duration, causality, the technology's contribution, and feasibility of mitigation. The baseline dimensions listed by Anthropic are:

  • Physical impacts: effects on bodily health and well-being
  • Psychological impacts: effects on mental health and cognitive functioning
  • Economic impacts: financial consequences and property considerations
  • Societal impacts: effects on communities, institutions, and shared systems
  • Individual autonomy impacts: effects on personal decision-making and freedoms

Tools and practices for mitigation

Depending on the type and severity of potential harm, Anthropic describes a range of policies and practices it uses to manage risks. These include maintaining a Usage Policy, carrying out evaluations before and after launches (including red teaming and adversarial testing), employing advanced detection techniques to spot misuse and abuse, and applying robust enforcement measures from prompt modifications to account blocking. The stated goal is to apply proportionate safeguards while preserving the everyday usefulness of the systems.

Examples of how the framework informs decisions

  • Computer use: As models gain the ability to interact with computer interfaces, Anthropic evaluates which kinds of software and contexts are involved. The company pays special attention to financial software and banking platforms—where unauthorized automation could enable fraud or manipulation—and to communication tools that could be exploited for targeted influence operations or phishing. Based on that analysis, Anthropic develops approaches meant to retain the utility of these capabilities while adding monitoring and enforcement to prevent misuse. Initial measures included raising enforcement thresholds and using hierarchical summarization to detect harms while maintaining privacy standards.

  • Model response boundaries: Anthropic describes the trade-offs between model helpfulness and appropriate limits. More helpful models can more easily provide information that violates Acceptable Use Policies or enables dangerous actions, while overly conservative models can refuse benign requests unnecessarily. In the case of Claude 3.7 Sonnet, Anthropic evaluated different request types along this spectrum and changed how the model handles ambiguous prompts to encourage safe, useful responses rather than blanket refusals. That work led to a 45% reduction in unnecessary refusals while keeping strong protections against genuinely harmful content. The company stresses that these trade-offs are especially important for vulnerable populations, such as children, marginalized communities, or people in crisis.

Next steps and collaboration

Anthropic acknowledges substantial further work remains. The company plans to continue evolving the framework, refine assessment methods, and learn from both successes and failures. It invites researchers, policy experts, and industry partners to collaborate and provides an email contact for engagement: usersafety@anthropic.com.

Anthropic notes the presented approach reflects its current thinking and will continue to develop as the company and the broader field gain more experience.