Business

AI-generated text

Anthropic raises $580M to develop steerable, interpretable and robust large language models

Anthropic announced a $580 million Series B round to fund research and infrastructure for making large language models more steerable, interpretable and robust.

Anthropic raises $580M to develop steerable, interpretable and robust large language models

Anthropic, an AI safety and research company founded in early 2021, has raised $580 million in a Series B financing round. The funds will be used to build large-scale experimental infrastructure and to study and improve the safety properties of computationally intensive AI models.

How the funding will be used

The company intends to develop technical components that give large language models better implicit safeguards and reduce the need for after-training interventions. It is also building tools to inspect models more deeply to validate that those safeguards work in practice. Anthropic plans to expand teams and partnerships tasked with exploring the policy and societal impacts of advanced models.

Research progress to date

Since its founding, Anthropic has focused on making systems more steerable, robust, and interpretable. Notable results cited by the company include:

  • Mathematical reverse-engineering of the behavior of small language models.
  • Initial insights into the sources of pattern-matching behavior in large language models.
  • Development of baseline techniques to make large language models more “helpful and harmless,” followed by reinforcement learning approaches to further improve these properties.
  • Release of a dataset intended to help other research labs train models that better align with human preferences.
  • An analysis of sudden performance changes in large language models and the societal impacts of this phenomenon, emphasizing the need to study safety issues at scale.

Leadership comments and organizational status

Dario Amodei, co-founder and CEO of Anthropic, said: “With this fundraise, we’re going to explore the predictable scaling properties of machine learning systems, while closely examining the unpredictable ways in which capabilities and safety issues can emerge at-scale.” He added that the company has made "strong initial progress on understanding and steering the behavior of AI systems," and is assembling the pieces needed to build usable, integrated AI systems that benefit society.

Daniela Amodei, co-founder and President, noted: “Now that we’ve built out the organization, we’re focusing on ensuring Anthropic has the culture and governance to continue to responsibly explore and develop safe AI systems as we scale.” Anthropic is now a growing team of around 40 people based in a plant-filled office in San Francisco, with plans to expand further this year.

Funding details

The Series B follows a $124 million Series A raised in 2021. The Series B round was led by Sam Bankman-Fried, CEO of FTX, and included participation from Caroline Ellison, Jim McClave, Nishad Singh, Jaan Tallinn, and the Center for Emerging Risk Research (CERR).

Why this matters

The substantial investment will allow Anthropic to scale computational and experimental capacity needed to probe and improve the safety behaviors of large models. By focusing on steerability, interpretability and robustness, and by studying societal impacts, the company aims to develop models that are easier to govern and better aligned with human needs.