Safety

OpenAI: accidental Chain of Thought evaluations occurred during training, monitoring not compromised

OpenAI recently built a system that scans all reinforcement learning (RL) runs, and during such checks they found some accidental Chain of Thought (CoT) evaluations during the training of previously deployed models.

OpenAI recently built a system that scans all reinforcement learning (RL) runs, and during such checks they found some accidental Chain of Thought (CoT) evaluations during the training of previously deployed models. The company says there is no clear evidence these degraded CoT monitorability, so model oversight and security are not expected to have been compromised.