Safety

AI-generated text

Anthropic recommends cybersecurity standards and two‑party controls for frontier AI models

Anthropic outlines cybersecurity practices it is adopting to protect frontier AI models and research, arguing that advanced AI requires protections beyond typical commercial standards.

Anthropic recommends cybersecurity standards and two‑party controls for frontier AI models

Date: July 25, 2023

As capabilities of frontier artificial intelligence models grow rapidly, Anthropic emphasizes that securing these systems is a critical priority. Building on previous posts about Anthropic’s safety approach and Claude’s capabilities, the company outlines steps it is taking and recommends industry and government actions to ensure models are developed and deployed securely.

Why this matters

Anthropic argues that advanced AI models could significantly alter economic and national security dynamics. Because the technology is strategically important, frontier AI research and models should be protected at levels well above typical commercial practices to prevent theft or misuse.

In the near term, governments and frontier AI labs must be prepared to protect advanced models, model weights, and the underlying research. Anthropic recommends developing robust, widely shared best practices and considering treatment of the advanced AI sector akin to “critical infrastructure” to enable deeper public–private cooperation in securing models and the companies that build them.

Cybersecurity best practices

Anthropic identifies “two‑party control” as necessary to secure advanced AI systems. This pattern — requiring more than one person to authorize sensitive actions — is already used in many domains (for example, the most secure vaults require two people with two keys) and appears in industry standards and review patterns across manufacturing (GMP, ISO 9001), food safety (FSMA PCQI, ISO 22000), medical devices (ISO 13485) and financial technology (SOX).

The company argues this approach should apply to every system involved in development, training, hosting, and deployment of frontier AI models. Practically, this means no individual should have persistent access to production‑critical environments; instead, time‑limited access must be requested from a colleague with a stated business justification. Anthropic calls this design “multi‑party authorization to AI‑critical infrastructure design” and describes it as a leading security requirement that depends on the full range of cybersecurity best practices for correct implementation.

Anthropic also highlights secure software development practices as essential to the frontier AI environment. It points to the NIST Secure Software Development Framework (SSDF) and the Supply Chain Levels for Software Artifacts (SLSA) as gold standards. The company notes that executive actions—such as the 2021 Executive Order 14028, which directed the Office of Management and Budget (OMB) to set federal procurement guidelines—can be effective levers to raise industry standards; EO 14028 motivated significant industry investment to meet SSDF‑like requirements to retain federal contracts.

Many aspects of SSDF and SLSA are transferable to model and model‑coupled software development. Anthropic argues that producing and deploying a model is analogous to building and deploying software, and that together SSDF and SLSA create a chain of custody: when properly applied, these practices help tie a deployed model back to the company that developed it, improving provenance.

Anthropic refers to this concept as a “secure model development framework” and urges extending SSDF to explicitly encompass model development within NIST’s standards process. In the near term, these practices could be enforced via procurement requirements on AI companies and cloud providers contracting with governments, which—given the role of U.S. cloud providers in hosting many frontier models—would have a market‑shaping effect even ahead of formal regulation.

Implementation and public‑private cooperation

Anthropic states it is implementing two‑party controls, SSDF, SLSA, and other cybersecurity best practices internally. The company acknowledges that as model capabilities scale, security protections will need further enhancement and that this will be an iterative process in consultation with government and industry.

Anthropic also recommends that frontier AI labs engage in public‑private cooperation similar to companies in other critical infrastructure sectors such as financial services. It suggests possibly designating the frontier AI sector as a special sub‑sector within the existing IT sector to facilitate enhanced information sharing and cooperation between labs and government agencies, helping defend against well‑resourced malicious cyber actors.

Conclusion

Anthropic warns against deprioritizing security for the sake of productivity: although security measures can sometimes impede workflows, there are creative ways to limit friction so research and operations can continue effectively. The company reiterates that AI development offers significant benefits but carries risks if not pursued thoughtfully. As the developer of Claude, Anthropic says it takes responsibility for building and deploying safe, secure, and human‑aligned systems and will continue to share perspectives on responsible AI development.