Safety

AI-generated text

Anthropic says Claude models are highly resistant to prompt injection

Anthropic announced that they have successfully trained their models to resist prompt-injection attacks that previously fooled early versions of Claude; similar results were observed in a benchmark…

Anthropic says Claude models are highly resistant to prompt injection

Anthropic announced that they have successfully trained their models to resist prompt-injection attacks that previously fooled early versions of Claude; similar results were observed in a benchmark and red teaming conducted by an independent researcher, which could reduce the risk of data exfiltration by agents.