Researchers presenting at a major machine-learning conference argue that language models cannot be made fully secure, because they judge where an instruction came from by its wording and style rather than by the labels that mark user input, system rules, outside documents or the model's own notes. Text written to imitate a model's private reasoning made several widely used models give banned answers, including drug synthesis and aircraft sabotage. The authors note the models tested were released last year, and one outside expert says leading models are now much harder to attack this way.
What changed
Defences against jailbreaks and prompt injection assumed models could be trained to tell apart instructions coming from a user, a system prompt, an outside document or their own notes.
What it unlocks
A concrete explanation of why role-based safety training keeps failing, and a test method: swap the role labels around text and see whether the model's behaviour changes at all.
- paper presented at ICML in July 2026
- attack won OpenAI's red-teaming hackathon in August 2025
- similar results reported on models from OpenAI, Anthropic, Alibaba and DeepSeek
- technologyreview.com2026-07-30