Safety

AI-generated text

Anthropic researcher warns >10% chance that future self-improving superintelligence could threaten humanity

An employee involved in training models at Anthropic resigned, arguing the company is not adequately addressing the risk that a future self-improving artificial superintelligence could threaten human survival.

Anthropic researcher warns >10% chance that future self-improving superintelligence could threaten humanity

According to the American Forbes, an employee involved in training AI models at Anthropic resigned, arguing the company is not adequately addressing the risk that a developed superintelligence could cause human extinction. The departing researcher, Jacob Coxon, previously worked at OpenAI.

What happened?

Jacob Coxon said that Anthropic is working on creating an artificial superintelligence that could potentially improve itself. He warned that such a self-improving system could have catastrophic consequences if its behavior cannot be safely constrained or aligned with human values.

Evan Hubinger, head of Anthropic's alignment research (the group focused on aligning AI with human values), responded to Coxon's post. Hubinger acknowledged that there is a possibility a powerful AI could cause the extinction of humanity. Based on his own estimate, he put the probability of that outcome at above 10 percent within the next ten years.

Which AIs are the concern?

Both Hubinger and Coxon emphasized that current models are not the primary danger. The central worry is the potential development over the next decade of a superintelligent system that can self-improve and thus surpass human capabilities. Hubinger stated that Anthropic currently lacks a ready, reliable solution for safely aligning such a system with human goals and values, and it is not clear whether that problem will be solved in time.

Broader debate and calls for slowing development

The resignation and internal debate reflect a wider conversation in the field. In July, several leading AI researchers — including Anthropic's co-founders and heads of other major AI labs — signed a statement saying there may be a need for industry and governments to deliberately slow the development of the most advanced AI systems. The rationale is to allow time to develop appropriate safety and regulatory frameworks.

Why this matters

The core issue is that future self-improving systems could present risks that go beyond current technical challenges. Decisions by researchers and labs about development pace and prioritization of safety work will materially affect how those risks evolve in the coming years.