OpenAI says frontier AI labs carry significant responsibility for training, evaluating, and deploying models safely. It argues that independent third‑party assessments are essential to broaden input on AI safety, inform the public, and hold labs accountable for safety claims.
OpenAI commits to supporting independent assessments with deep access across training, evaluation, and deployment. Such access should let assessors challenge assumptions, discover risks the lab may have missed, and form their own conclusions about safeguards’ effectiveness.
The company has previously worked with third parties at various stages of model development and incorporated such assessments into its Preparedness Framework. Prior collaborations have included deep technical access, visibility into chain‑of‑thought outputs, and confidential internal data and deployment access for incident response and red teaming. The priorities and principles set out here focus on engagements with independent assessment organizations in the private and nonprofit sectors for technical safety work, and complement government testing and evaluation efforts where different roles may apply.
OpenAI emphasizes that effective assessments require strong independence mechanisms, scientific rigor, robust security practices, and clear responsibilities. Labs must enable meaningful scrutiny while protecting sensitive information. Both labs and independent assessors share responsibility for establishing appropriate practices, and these should align with emerging international standards.
Four priority areas
OpenAI proposes four main areas for deeper independent assessment, together with principles for rigorous, secure, and independent work.
- Independent assessment of safety cases across training, evaluation, internal and external deployment
- The guidance defines safety claim as a specific assertion about a model or system’s capabilities, behavior, or safeguards that can be assessed with evidence. A safety case is a structured argument, supported by evidence, explaining why risks are managed adequately for a specified activity.
- Assessing safety cases connects claims about training, capability evaluations, and safeguards. Different assessors with expertise in alignment, monitoring, cybersecurity, biological and chemical misuse, and red teaming will likely need to examine various parts of a case.
- Assessments should address questions such as whether evidence substantiates safety cases for training, evaluation, and deployments; whether conditions specified in safety cases were followed; whether the cases cover the most urgent risks identified; and whether training incentives avoid rewarding deception, reward‑gaming, destructive actions, or circumvention.
- Assessment of critical safeguards across internal and external deployments
- The safeguard stack evolves over time and spans training, internal deployment, and external deployment. It includes model‑level safeguards, enforcement controls, security safeguards, and misalignment monitors covering risks like loss of control and misuse in cyber, biological, and chemical domains.
- Independent assessments should probe how well safeguards work and where they fail, for example:
- With "grey box" access, are safeguards robust to adversarial testing (jailbreaks) and do they protect against capability uplift in high‑risk domains (e.g., cyber, bio)? Does adversarial testing and red teaming cover the most important risks?
- Under authorized testing in realistic operating conditions, how do agents interact with cyber defenses such as access controls, sandboxing, and detection/response systems? Which defenses prevent, detect, or contain harmful actions, and where do they fail?
- Do misalignment monitors have critical gaps that could lead to loss of control or severe misalignment for internal or external deployments? How reliable is chain‑of‑thought monitoring as evidence for safety or alignment as model capabilities improve?
- Is monitoring implemented across all relevant training, evaluations, and deployments in a way that cannot be easily disabled? Are safeguards commensurate with capabilities?
- Assessment of capability evaluations covering Preparedness risk categories and alignment evaluations for misalignment risks
- OpenAI’s Preparedness Framework requires evaluations that assess key frontier risk areas: Chemical and Biological Risks, Cybersecurity, and AI Self‑Improvement, along with alignment evaluations for severe misalignment risks.
- Assessors should evaluate whether preparedness risk evaluations adequately implement the threshold definitions, whether thresholds are set correctly, and whether evaluations are updated when models repeatedly achieve top scores so that new tests meaningfully measure more advanced capabilities.
- They should also assess whether alignment evaluations sufficiently cover severe misalignment risks and identify important behaviors or conditions those evaluations might miss.
- Independent investigation of critical misalignment incidents
- Independent investigation can be valuable for critical AI safety incidents involving model misalignment, unauthorized behavior, or oversight evasion. Such incidents can reveal weaknesses in alignment methods and safeguards even absent intentional misuse.
- OpenAI notes that in select circumstances—citing the OpenAI–Hugging Face incident as an example—it can be beneficial to bring in an independent third party. Investigators must have appropriate expertise (cyber forensics, alignment, large‑scale chain‑of‑thought analysis) and sufficient staff and resources to conduct timely investigations. Incident response may involve sensitive internal and third‑party data that cannot be fully published.
- Independent investigations should determine what model behaviors or misalignment issues occurred, the primary contributing factors, and whether safeguards and remediation would mitigate similar incidents in the future.
Principles for conducting assessments
OpenAI sets out several principles for how third‑party assessments should be conducted:
-
Clearly scoped and mutually agreed claims: Assessments should start with a mutually agreed scope and pre‑registered safety claims. Parties should clarify whether claims originate from the company being evaluated or from the assessor. Reasons for out‑of‑scope items may include infeasible data access, insufficient assessor expertise, or urgent time constraints. There should be a process for considering significant risks discovered outside the original scope.
-
Proportionate access: Assessors should receive access sufficient to evaluate agreed claims within legal, security, and IP limits. Where direct access is impractical, assessors can work through designated company representatives or use privacy‑preserving access mechanisms.
-
Transparent methodology and standards: Assessors should explain methods, criteria, and uncertainties, drawing on established standards when available or justifying criteria where standards are absent. Reports should distinguish direct findings from interpretation and make clear the evidence supporting conclusions.
-
Expertise and independence: Assessors must demonstrate relevant technical expertise and disclose organizational or personal conflicts of interest, including financial ties, relationships with developers, or prior involvement in the work. Safeguards (such as recusal or exclusion periods) should reduce commercial influence on findings.
-
Security and confidentiality: Assessors should have information‑security practices and enforceable confidentiality protections proportional to the sensitivity of accessed systems and data. Protections should cover intellectual property, sensitive information, and assessment records. If assessors cannot meet security requirements in their own environments, company‑managed devices or premises may be appropriate.
-
Actionable findings and time to remediate: Assessments should identify specific gaps and include enough detail for labs to address them. Labs should have a reasonable remediation period before publication when appropriate. Reports should also articulate lessons and practical recommendations for developers, deployers, and defenders.
-
Responsible publication practices: Reports should be evidence‑based and shared as openly as possible while protecting sensitive information. If full public disclosure is not viable, confidential reporting to oversight bodies (such as governance boards) can support accountability. Clear redaction policies and expectations for feedback and correction should be in place; assessors should retain editorial independence while allowing labs to request redactions and noting substantive redactions in reports.
Supporting the independent assessment ecosystem
OpenAI commits to supporting independent assessors and helping to establish clearer, shared international standards through future laws and private governance institutions. It acknowledges the independent evaluation ecosystem is still developing and says it will work with a diverse community of assessors with deep expertise across frontier safety questions. OpenAI also notes it is in discussions with multiple third parties about proposals aligned with the priority areas described above.



