A fundamental flaw leaves LLMs strikingly vulnerable to attack
Researchers at ICML argue LLMs can't be fully secured because they rely on role tags to tell who said what, and attackers can forge those tags. Using 'chain-of-thought forgery,' they got OpenAI's gpt-oss-20b and GPT-5 to output instructions for making cocaine and sabotaging aircraft navigation. The team says red-teaming and patching can't fix this—it's a structural dead end. Similar results were seen on models from Anthropic, Alibaba, and DeepSeek, though the post doesn't name specific models or share test details.
Why it matters: ICML paper reveals a 'chain-of-thought forgery' attack targeting role tags, tested successfully against two OpenAI models. Concrete method + named targets make it solid. Held back from 85 because it's a conference report without a full paper or patch yet.