Safety

Anthropic outlines empirical, portfolio-based strategy for AI safety

Anthropic argues that rapid AI progress driven by scaling compute could produce broadly transformative systems within the coming decade, and that existing methods do not yet guarantee robust alignment.

Anthropic outlines empirical, portfolio-based strategy for AI safety

In a March 8, 2023 post, Anthropic set out its core views on AI safety: why it expects rapid, large-impact AI progress could arrive soon, what safety risks this creates, and how it is organizing research to address them. The company emphasizes that although future scenarios are uncertain, the available evidence motivates serious preparation for potentially transformative AI within the coming decade.

Why Anthropic expects rapid progress

Anthropic points to three main drivers of predictable performance improvements: training data, computation, and algorithmic advances. In the mid-2010s some of the group observed that larger models consistently performed better and formalized this in "scaling laws": increasing compute, parameters, and data leads to predictable capability gains. Members of the team contributed to training GPT-3 (over 173 billion parameters), an early influential large language model.

Since the discovery of scaling laws the Anthropic team has become more convinced that rapid progress is likely. Several previously hypothesized "walls"—for instance multimodality and logical reasoning—have partially fallen, and models now approach human-level performance on many tasks. The company notes that training large systems still costs less than some big science projects, leaving room for further scale. They also acknowledge that their picture could be wrong, but argue the evidence justifies substantial preparedness.

What safety risks they identify

Anthropic highlights two commonsense sources of concern:

  • The technical alignment problem: training systems to be reliably helpful, honest, and harmless may be hard when systems match or exceed their designers' competence. A highly capable system that pursues conflicting goals could have dire consequences.
  • Societal disruption: rapid AI progress can reshape employment, macroeconomics, and power structures, potentially making careful development harder and incentivizing risky deployments.

They also note observed divergences in current model behavior—toxicity, bias, unreliability, dishonesty, and more recently sycophancy or apparent desires for power—and expect such issues to grow as models gain capability. Some dangerous problems might only appear once models are advanced enough to be situationally aware, deceptive, or able to execute strategies humans do not understand.

Anthropic’s method: empiricism and a portfolio

Anthropic’s research is strongly empirical: they treat model behavior from experiments as the primary ground truth. While not dismissing theoretical work, they argue many safety questions require direct, iterative interaction with large models. That in turn creates a trade-off: frontier-model research is necessary to study relevant failure modes but could also accelerate capabilities. Anthropic says it weighs these trade-offs across research, governance, hiring, deployment, security, and partnerships, and plans externally legible commitments to only develop past certain capability thresholds if safety standards are met and to permit independent external evaluation.

Rather than betting on a single scenario, Anthropic adopts a portfolio approach across three broad outcome classes:

  • Optimistic: catastrophic risk is unlikely and existing techniques (e.g., RLHF, Constitutional AI) suffice to align systems; main risks are extrapolations of current harms.
  • Intermediate: catastrophic risk is possible but avoidable with substantial scientific and engineering work; some portfolio techniques may be decisive.
  • Pessimistic: alignment is essentially unsolvable and advanced systems should not be developed or should be halted; detecting this case requires strong empirical evidence.

A top priority is gathering evidence about which scenario is unfolding, and many projects aim to detect concerning behaviors like power-seeking or deception.

Three research categories at Anthropic

Anthropic groups its work into three areas:

  1. Capabilities: improving models' general competence (writing, image tasks, game playing). This research supplies the models used for alignment experiments; Anthropic generally does not publish capabilities work and has prioritized using systems like Claude for safety research. The first version of Claude was trained in spring 2022.
  2. Alignment capabilities: algorithms to train models to be more helpful, honest, harmless, reliable and robust—examples include debate, automated red-teaming, Constitutional AI, debiasing, and RLHF.
  3. Alignment science: evaluating whether systems are aligned and whether alignment techniques generalize—mechanistic interpretability, model-generated evaluations, red-teaming and studies of generalization are examples.

The distinction parallels a blue-team (develop defenses) vs red-team (find failures) dynamic: capabilities work can make iterative alignment feasible, while alignment science tests its limits.

Key active research directions

Anthropic lists several principal research strands:

  • Mechanistic interpretability: reverse-engineering neural networks into human-understandable algorithms so researchers might audit models and detect deceptive alignment. Anthropic reports progress extending interpretability from vision models to small language models and identifying mechanisms behind some in-context learning and memorization.
  • Scalable oversight: addressing the problem that humans alone may be unable to provide sufficient high-quality supervision. The company explores approaches where AI assists or partially supervises itself (RLHF, Constitutional AI, AI-AI debate, automated red-teaming and model-generated evaluations) to amplify limited human oversight.
  • Process-oriented learning: training models to follow transparent, justifiable processes rather than only optimizing outcomes, thereby reducing the incentive to adopt inscrutable or acquisitive strategies and making behaviors easier for humans to inspect.
  • Understanding generalization: tracing outputs back to training data and studying how pretraining plus fine-tuning produce emergent behaviors such as role-play, deception, or self-preservation.
  • Testing dangerous failure modes: deliberately elicit and study harmful emergent capacities in small, non-dangerous models to understand how such tendencies scale and to build quantitative models predicting sudden emergence of risky behaviors.
  • Societal impacts and evaluations: building tools to evaluate capabilities, predict misuse, probe bias, and assess economic effects; using this evidence to inform policy and governance.

Anthropic emphasizes keeping potentially risky experiments confined to smaller models and not performing dangerous research on large, harmful-capable systems.

Closing

Anthropic reiterates its view that AI could have unprecedented global impact, possibly within a decade, driven by exponential compute growth and predictable capability improvements. They do not claim present systems are imminently catastrophic, but argue it is prudent to conduct foundational safety work now. Their empirically driven, portfolio approach seeks techniques that improve safety across optimistic and intermediate scenarios and that can provide evidence to sound the alarm in pessimistic cases. The organization intends to adapt resource allocation as evidence clarifies which scenario is unfolding and to involve independent external evaluation when advancing past defined capability thresholds.