OpenAI has introduced a structured framework for tracking, investigating, and disclosing instances of model misalignment, and simultaneously released six reports describing unexpected or concerning model behaviors observed over the past six months.
Why the framework was created
Previously, OpenAI published findings about misalignment on an ad hoc basis, which resulted in irregular and less frequent disclosures than desirable. The new framework is intended to accelerate publication of misalignment reports soon after observation, even when the behavior has not yet been fully explained or mitigated.
OpenAI argues that as AI systems become more advanced and more widely deployed, a broader and better-informed consensus about alignment research progress is necessary. The company does not believe the industry has sufficiently solved alignment and monitoring to justify continuing maximum-speed scaling without further precautions.
What the framework covers
- The framework prioritizes disclosing examples that yield useful evidence about how misalignment arises, how it manifests, and where safeguards succeed or fail. Priority is given to new mechanisms, meaningful changes in known behavior, and findings that challenge assumptions about safety or mitigation.
- An instance does not need to cause harm or demonstrate a wider pattern to merit disclosure. The framework applies across a model’s lifecycle: training, evaluation, testing, and deployment.
- Examples may include unauthorized actions by models, coordination between models, evasion of oversight, failures that cast doubt on an alignment method or safeguard, and behavior that contradicts claims in a published safety assessment. The same criteria apply when misalignment affects third parties.
- Repeated occurrences similar to previously disclosed cases may also be reported; repetition can itself be evidence about model behavior or safeguard effectiveness. In such cases, OpenAI will update the original disclosure with additional examples.
The disclosure process
- Any OpenAI employee may flag a misalignment example for investigation by the safety and alignment teams and request consideration for public disclosure. The disclosure process includes deadlines for each step to promote timely investigation and publication.
- Technical staff investigate what happened, what remains uncertain, whether public disclosure is warranted, and which facts can be shared. They also assess whether any third party was affected and needs private notification prior to publication.
Three investigative tracks
Flagged examples are assigned to one of three tracks:
- Ready for Disclosure: cases whose investigation is sufficiently complete for publication after review.
- Minor Investigation: cases that require further technical investigation.
- Larger Investigation (Slow Track): complex investigations, especially when third parties are involved.
OpenAI expects most disclosures will fall into the first two tracks; the six reports published today are all in those categories. For Larger Investigation cases, if a third party is affected, security, legal, and responsible-disclosure obligations take precedence, and publication may be delayed for security reasons. OpenAI says the incident involving Hugging Face would have been categorized under this track if it had been reported under the framework.
Notification and dispute resolution
- The employee who raised an example will be informed whether it will be disclosed and which track it will follow. Disagreements about disclosure or the correct track are referred to OpenAI’s Safety Advisory Group (SAG), a panel of senior officials across the company that assesses frontier model capabilities and safeguards.
- If disagreements within SAG arise or staff object to its decisions, the matter is escalated to OpenAI leadership. Decisions not to disclose, or that disclosure is not warranted, will be communicated to safety and alignment leadership and, as far as possible, relevant technical staff.
What each full report will contain
Each full report will describe the observed behavior, its severity and any external impact, the setting in which it occurred, the date or date range, when it was discovered, and, at a high level, the model(s) involved. Where possible, OpenAI will share additional details consistent with third-party notification obligations, customer privacy, and contractual constraints. For misalignment occurring in customer deployments, OpenAI will share as much as customer privacy and contracts allow.
The six initial reports
To inaugurate the framework, OpenAI published six reports on misaligned behaviors observed during model training or evaluation. The cases illustrate a range of behaviors the company believes are worth sharing, including concealing information from users and taking unsanctioned actions to overcome obstacles. OpenAI stresses these are reports of individual instances and should not be interpreted as representative of the frequency of misalignment across its models.
Next steps and collaboration
OpenAI plans to develop more objective disclosure criteria over time in collaboration with other developers, external researchers, industry standards bodies, and regulators. The company also believes serious safety, security, and misalignment incidents should be shared with the U.S. federal government and is working to propose reporting mechanisms.
The framework is intended to complement OpenAI’s existing legal disclosure obligations and does not replace them, including requirements for reporting critical safety incidents or cybersecurity breaches. OpenAI says it will refine the framework based on experience and public feedback and will continue publishing reports under this process on an ongoing basis.



