Why prompt injection keeps working: ICML paper traces it to role confusion in LLMs
A Theory of Why Prompt Injection Works
This ICML 2026 paper reframes prompt injection as role confusion. LLMs receive a single text stream with role tags (system, user, think, tool) to separate instructions from external data. Attackers hide commands inside tool-tagged webpage content; the model fails when it misreads tool text as a user instruction. The authors argue current models rely on attack memorization—scoring well on static benchmarks but easily bypassed by human red-teamers who rephrase attacks. The robust fix is role perception: respecting role tags instead of pattern-matching known attacks. The post does not disclose a concrete defense or deployment timeline.
Why it matters: ICML 2026 paper attributing prompt injection to role-tag perception flaws, with reproducible attacks and mechanistic explanations. All three HKR axes hit. Score held below 85 because it's a research paper, not a product update, and the accessibility bar is slightly high for no...