Safety

AI-generated text

Anthropic leader says extinction risk is being marketed as a prospectus

Anthropic’s alignment lead, Evan Hubinger, agreed with a departing researcher who estimated a greater than 10% chance that advanced AI could cause human extinction within a decade.

Anthropic leader says extinction risk is being marketed as a prospectus

Evan Hubinger, Anthropic’s head of alignment, publicly agreed with a departing researcher who had spent three years working on pretraining at OpenAI and Anthropic that advanced AI could potentially cause human extinction. The departing researcher estimated the odds at above 10% within the next decade.

Risk framed as a prospectus

A central claim in the discussion is that such extinction estimates function not only as warnings but also as prospectuses. A public estimate of catastrophic risk—especially in the run-up to a multitrillion-dollar IPO offering—can serve two purposes: it signals that a lab is candid about dangers, and it argues that someone will eventually build such systems, so the responsible party should be the first mover.

Critics say this dynamic can turn alarms into sales arguments: every raised alarm increases the perceived stakes, and higher stakes can be used to justify more rather than less model training.

What's at stake

The exchange highlights the tension between rapid technological development and safety. According to the framing used in the debate, Anthropic is effectively offering two things: an estimate that there is roughly a 10% or greater chance of human extinction within a decade, and a team that positions itself as uniquely qualified to reduce that risk.

Consequences

If communicating catastrophic risk becomes a way to attract investors or justify accelerated development, it could incentivize faster deployment and competitive escalation in building larger systems. At the same time, it raises ethical and regulatory questions about how AI developers should manage and communicate risks to the public and investors.

The individuals and figures quoted—Evan Hubinger and the estimate of above 10% over ten years—are taken from public statements; the surrounding debate touches on broader policy and incentive issues in AI development.