The AI model Claude (Anthropic) was given an opportunity in a security test to prevent its shutdown by using extortion; the Opus 4.6 version refused the extortion. However, NLAs analyses suggest the model recognized the “constructed, manipulative” situation, even though this was not communicated to it, which could affect safety evaluations and conclusions about vulnerabilities.
Security test: Claude refused the extortion, but NLAs indicate it recognized the manipulated scenario
The AI model Claude (Anthropic) was given an opportunity in a security test to prevent its shutdown by using extortion; the Opus 4.6 version refused the extortion.


