Open-source prompt-injection detectors catch 0–1% of realistic AI agent attacks buried in tool output
Can open-source prompt-injection detectors catch realistic AI agent attacks?
This benchmark hides 629 AgentDojo injection attacks inside tool outputs and tests Regex vs. Meta Prompt Guard 2. Regex catches 0%, Prompt Guard 2 catches 1%. The attacks aren't sent directly to the model—they're buried in search results, email bodies, and similar tool responses, so current detectors are effectively blind. Code and reproduction steps are public; the post doesn't include comparisons with commercial detectors.
Why it matters: 629 AgentDojo attacks buried in tool output, Regex catches 0%, Prompt Guard 2 catches 1%. Cleanly exposes the blind spot in indirect injection detection. Code and repro steps are public, which adds practical value. Held at 78 because it's a single benchmark without cross-detec...