Dario Amodei, CEO of Anthropic, presented the company’s Responsible Scaling Policy (RSP) and the related AI Safety Levels (ASL) framework in prepared remarks at the AI Safety Summit. The remarks were published by Anthropic on November 1, 2023. Anthropic originally published its RSP in September and was the first major AI company to make such a policy public.
Why an RSP?
Amodei emphasized the rapid pace of AI progress: systems that a few years ago struggled to form coherent sentences can now pass medical exams and perform other sophisticated tasks. He attributed much of this acceleration to the growth in available compute, which he said is increasing roughly eightfold per year. While broad trends are predictable, it is difficult to forecast when specific capabilities will appear — including dangerous ones such as the ability to provide information that could enable biological weapons.
The Responsible Scaling Policy aims to do two things: provide frequent, regular testing for dangerous capabilities as models are scaled, and establish a protocol for how to respond if such capabilities are detected. The idea of responsible scaling was initially suggested by the Alignment Research Center.
The AI Safety Levels (ASL) system
Anthropic designed the ASL system loosely after the internationally recognized biosafety level (BSL) system. Each ASL level uses an if–then structure: if a model exhibits certain dangerous capabilities, then deployment or further training is disallowed until specified safeguards are in place.
- ASL-1: Models with little or negligible risk, for example a specialized chess-playing AI.
- ASL-2: Describes the current era — models that present a range of present-day risks but do not yet have capabilities that would enable catastrophic misuse in domains like biology or chemistry. For ASL-2 models the RSP requires current best practices, including model cards, external red-teaming, and strong security.
- ASL-3: Triggered when models become operationally useful for catastrophic misuse in CBRN (chemical, biological, radiological, nuclear) areas, as judged by domain experts and in comparison to existing capabilities and proofs of concept. When ASL-3 is reached Anthropic requires:
- Unusually strong security such that non-state actors cannot steal model weights and state actors would need to expend significant effort to obtain them.
- Deployed versions of ASL-3 models must never produce information that would operationally increase CBRN risks, even when red-teamed by world experts working alongside AI engineers. Achieving this will likely require research breakthroughs, but Anthropic considers it a necessary safety condition.
- ASL-4 must be rigorously defined before ASL-3 is reached.
- ASL-4: Represents an escalation beyond ASL-3 in catastrophic misuse risk and introduces a new risk: autonomous AI systems escaping human control and posing significant societal threats. ASL-4 would be triggered roughly when systems achieve near-human levels of autonomy or become the primary source worldwide of at least one serious global security threat (for example, bioweapons). At ASL-4, Anthropic expects to require a detailed and precise understanding of what is happening inside models in order to make an affirmative case that a model is safe.
Operational lessons and company commitments
Amodei stressed that deep executive involvement matters: he personally spent 10–20% of his time on the RSP for three months, writing multiple drafts and proposing the ASL system. One co-founder devoted 50% of their time to the RSP development for three months. These leadership investments were intended to signal to employees that the company is seriously committed to safety and responsible scaling.
Anthropic recommends turning RSP protocols into product and research requirements so they drive planning and team roadmaps. At Anthropic this meant ramping up hiring in security, trust and safety, red teaming, and interpretability teams to have a reasonable chance of meeting ASL-3 safety measures by the time ASL-3 models appear.
Accountability is also required. Anthropic’s RSP is a formal directive of its board and is ultimately accountable to the Long Term Benefit Trust, an external panel of experts with no financial stake in Anthropic. Operationally, the company will implement a whistleblower policy before reaching ASL-3 and has already appointed an officer responsible for RSP compliance who reports to the Long Term Benefit Trust. As risks increase, Anthropic expects stronger forms of accountability will be necessary.
RSPs and regulation
Amodei made clear that RSPs are not intended to replace regulation but can serve as prototypes for it. He did not suggest that Anthropic’s RSP should be literally codified into law; rather, the RSP is a first attempt to address a difficult problem and is likely imperfect. As companies implement and iterate on RSPs, they will learn how to operationalize commitments in practice. Anthropic hopes RSP-style frameworks will be refined across companies and that governments can take the best elements to craft testing and auditing regimes with accountability and oversight, creating a "race to the top" where actors build on each other’s ideas to manage AI risks without unnecessarily blocking benefits.
Summary
Anthropic’s RSP and ASL framework aim to detect dangerous capabilities early as models scale and to define concrete, enforceable requirements for further training and deployment. The company has embedded RSP obligations into governance and operational planning, tied them to increased hiring in relevant safety teams, and set up external oversight through the Long Term Benefit Trust, with further accountability measures to follow as risks grow.



