Safety

Role confusion and “destyling”: how writing style enables prompt injection
Safety

Role confusion and “destyling”: how writing style enables prompt injection

Researchers Charles Ye, Jasmine Cui, and Dylan Hadfield-Menell show that large language models can be misled by the style of text that mimics internal role tags (e.g., <system>, <think>, <assistant>), a phenomenon they name “role confusion.” They demonstrate a mitigation called “destyling” — rewriting attacker text to look less like model-internal blocks — which reduced attack success in their dataset from 61% to 10%.

3 min read