Skip to content
Hacker News front page

Why prompt injection keeps working: ICML paper traces it to role confusion in LLMs

A Theory of Why Prompt Injection Works

This ICML 2026 paper reframes prompt injection as role confusion. LLMs receive a single text stream with role tags (system, user, think, tool) to separate instructions from external data. Attackers hide commands inside tool-tagged webpage content; the model fails when it misreads tool text as a user instruction. The authors argue current models rely on attack memorization—scoring well on static benchmarks but easily bypassed by human red-teamers who rephrase attacks. The robust fix is role perception: respecting role tags instead of pattern-matching known attacks. The post does not disclose a concrete defense or deployment timeline.

Why it matters: ICML 2026 paper attributing prompt injection to role-tag perception flaws, with reproducible attacks and mechanistic explanations. All three HKR axes hit. Score held below 85 because it's a research paper, not a product update, and the accessibility bar is slightly high for no...

Read the original ↗Export Markdown