Regulation

Anthropic’s recommendations for building an AI accountability framework

On June 13, 2023, Anthropic submitted its response to the National Telecommunications and Information Administration’s (NTIA) Request for Comment on AI Accountability, outlining concrete proposals for evaluating and governing advanced AI systems.

Anthropic’s recommendations for building an AI accountability framework

On June 13, 2023, Anthropic submitted its response to the National Telecommunications and Information Administration’s (NTIA) Request for Comment on AI Accountability and published a set of recommendations that reflect the company’s core AI policy proposals.

Anthropic notes that there is currently no robust, comprehensive process for evaluating today’s advanced artificial intelligence (AI) systems, nor for the more capable systems that may appear in the future. Their submission outlines the processes and infrastructure they consider necessary to ensure AI accountability. It also discusses the NTIA’s potential role as a coordinating body that could set standards in collaboration with other government agencies such as the National Institute of Standards and Technology (NIST).

Key recommendations

  • Fund research to build better evaluations

    • Increase funding for AI model evaluation research. Developing rigorous, standardized evaluations is difficult and resource-intensive, and government funding could accelerate progress in this critical area.
    • In the near term, require companies deploying AI systems to disclose their evaluation methods and results. Disclosures need not be public if doing so would compromise intellectual property or confidential information, but transparency would help researchers and policymakers identify gaps in existing evaluations.
    • In the long term, develop industry evaluation standards and best practices that government agencies like NIST could help establish and coordinate.
  • Create risk‑responsive assessments based on model capabilities

    • Develop standard capability evaluations for AI systems, with a focus on critical risks from advanced AI such as deception and autonomy. Rigorous capability and safety evaluations can form an evidence‑based foundation for proportionate, risk‑responsive regulation.
    • Conduct more research and provide funding to establish a risk threshold. Once defined, mandate evaluations of all models against this threshold:
      • If a model falls below the threshold, existing safety standards are likely sufficient: verify compliance and permit deployment.
      • If a model exceeds the threshold and safety assessments or mitigations are insufficient, halt deployment, strengthen oversight, and notify regulators. Determine appropriate safeguards before permitting deployment.
  • Establish pre‑registration for large AI training runs

    • Create a process for AI developers to report large training runs so regulators are aware of potential risks. This requires deciding the appropriate recipient, required information, and cybersecurity, confidentiality, IP, and privacy safeguards.
    • Specifically, establish a confidential registry in which AI developers can pre‑register large training runs with their home country’s national government (for example, model specifications, model type, compute infrastructure, intended training completion date, and safety plans) before training begins. Aggregated registry data should be protected to the highest available standards.
  • Empower technically literate, security‑conscious, and flexible third‑party auditors

    • Third‑party auditors should meet three key criteria:
      • Technically literate — some auditors must have deep machine learning experience;
      • Security‑conscious — able to protect valuable IP that could pose national security risks if stolen;
      • Flexible — able to conduct robust but lightweight assessments that surface threats without undermining U.S. competitiveness.
  • Mandate external red teaming before model release

    • Require external red teaming for AI systems as a precondition for releasing advanced models, either through a centralized third party (e.g., NIST) or via decentralized mechanisms (e.g., researcher API access), to standardize adversarial testing.
    • Ensure high‑quality external red teaming options exist before making this a general requirement, since red teaming expertise currently resides mostly within private AI labs.
  • Advance interpretability research

    • Increase funding and provide government grants and incentives for interpretability research at universities, nonprofits, and companies, enabling meaningful work on smaller models and progress outside frontier labs.
    • Acknowledge that regulations demanding interpretable models would be infeasible today, but may become possible with future research advances.
  • Clarify antitrust rules to enable industry collaboration on safety

    • Regulators should issue guidance on permissible AI industry safety coordination under current antitrust laws. Reducing legal uncertainty would allow private companies to collaborate on public‑interest safety goals without risking antitrust violations.

Conclusion

Anthropic believes these recommendations would bring stakeholders closer to an effective AI accountability framework. The company emphasizes that implementation will require collaboration among researchers, AI labs, regulators, auditors, and other stakeholders, and states its commitment to supporting safe development and deployment of AI systems. Elements such as evaluations, red teaming, standards, interpretability and other safety research, auditing, and strong cybersecurity practices are highlighted as promising ways to mitigate AI risks while realizing benefits.

Anthropic made its full submission to the NTIA public in the official document accompanying their response.