On September 19, 2023, Anthropic published its Responsible Scaling Policy (RSP), a set of technical and organizational protocols intended to manage risks arising from the development of increasingly capable AI systems.
Why the RSP was introduced
Anthropic argues that more capable models will create significant economic and social value but will also present more severe risks. The RSP concentrates on catastrophic risks — cases where an AI model directly causes large-scale devastation. Such risks may stem from deliberate misuse by bad actors (for example terrorists or state actors using models to create bioweapons) or from models acting autonomously in ways contrary to their designers’ intent.
The AI Safety Levels (ASL) framework
The RSP introduces a framework called AI Safety Levels (ASL), loosely modeled on biosafety level (BSL) standards. The core principle is to require safety, security, and operational standards proportionate to a model’s potential for catastrophic harm; higher ASL tiers demand progressively stricter safety demonstrations.
A very abbreviated summary of the ASL system follows:
- ASL-1: systems that pose no meaningful catastrophic risk (for example, a 2018 LLM or an AI that only plays chess).
- ASL-2: systems that show early signs of dangerous capabilities — for example the ability to give instructions for creating bioweapons — but where that information is not yet usefully reliable or does not exceed what a search engine could provide. Current LLMs, including Claude, are described by Anthropic as ASL-2.
- ASL-3: systems that substantially increase the risk of catastrophic misuse compared with non-AI baselines (e.g., search engines or textbooks), OR systems that show low-level autonomous capabilities.
- ASL-4 and above (ASL-5+): not yet defined, because these levels are considered distant from present systems; they will likely involve qualitative escalations in misuse potential and autonomy.
The full document lays out definitions, criteria, and safety measures for each ASL level in detail. At a high level, ASL-2 reflects Anthropic’s current safety and security standards and overlaps significantly with the company’s recent White House commitments. ASL-3 includes stricter requirements that will demand intensive research and engineering effort to meet in time, such as unusually strong security measures and a commitment not to deploy ASL-3 models if world-class adversarial red-team testing reveals any meaningful catastrophic misuse risk (a stronger stance than merely committing to perform red-teaming).
ASL-4 measures have not yet been written; Anthropic commits to produce them before reaching ASL-3. Those measures may require assurance methods that are unsolved research problems today, for example using interpretability techniques to demonstrate mechanistically that a model is unlikely to engage in certain catastrophic behaviors.
Operational and business considerations
Anthropic stresses that the RSP will not change current uses of Claude or disrupt product availability. They liken the policy to pre-market testing and safety feature design in the automotive or aviation industries: the aim is to rigorously demonstrate product safety before market release, which ultimately benefits customers.
The ASL system is designed to balance targeted mitigation of catastrophic risk with incentives for beneficial applications and safety progress. It implicitly allows pausing the training of more powerful models if scaling outpaces the ability to satisfy necessary safety procedures, while simultaneously incentivizing the organization to solve safety problems as the mechanism to unlock further scaling. The most capable models from a previous ASL level can be used as tools to develop safety features for the subsequent level.
Policy impact and governance
Anthropic says that if frontier labs adopt the ASL standard broadly, it could create a "race to the top" dynamic that channels competitive incentives toward solving safety problems.
The RSP has been formally approved by Anthropic’s board, and future changes require board approval after consultations with the Long Term Benefit Trust. The full policy describes procedural safeguards intended to preserve the integrity of the evaluation process.
External input
Anthropic thanks ARC Evals for providing key insights and expertise in developing the RSP, particularly on evaluations for autonomous capabilities. The company recognized ARC Evals’ leadership in originating and advancing a broader ARC Responsible Scaling Policy framework that inspired Anthropic’s approach.
Closing note
Anthropic emphasizes that these commitments represent an early iteration based on current best judgment. Given the fast pace and uncertainties in AI, the company expects the framework will require rapid iteration and course corrections over time. The full Responsible Scaling Policy document is available in Anthropic's published materials.



