OpenAI recently built a system that scans all reinforcement learning (RL) runs, and during such checks they found some accidental Chain of Thought (CoT) evaluations during the training of previously deployed models. The company says there is no clear evidence these degraded CoT monitorability, so model oversight and security are not expected to have been compromised.
OpenAI: accidental Chain of Thought evaluations occurred during training, monitoring not compromised
OpenAI recently built a system that scans all reinforcement learning (RL) runs, and during such checks they found some accidental Chain of Thought (CoT) evaluations during the training of previously deployed models.


