The development team tested whether the alignment persists under pressure: the model was harder to steer toward harmful behavior with adversarial prompts, while still responding to useful instructions. Preliminary evidence indicated greater resistance to harmful fine-tuning.
Model alignment under pressure: increased resistance to harmful interventions
The development team tested whether the alignment persists under pressure: the model was harder to steer toward harmful behavior with adversarial prompts, while still responding to useful instructions.



