Anthropic announced that they have successfully trained their models to resist prompt-injection attacks that previously fooled early versions of Claude; similar results were observed in a benchmark and red teaming conducted by an independent researcher, which could reduce the risk of data exfiltration by agents.
AI-generated text
Anthropic says Claude models are highly resistant to prompt injection
Anthropic announced that they have successfully trained their models to resist prompt-injection attacks that previously fooled early versions of Claude; similar results were observed in a benchmark…



