Skip to content
Trending storyPast story

Open-source prompt-injection detectors catch 0–1% of realistic AI agent attacks buried in tool output

1 report1 sourceupdated 6 days ago

What happened

Summary

这个基准测试把 629 条来自 AgentDojo 的注入攻击藏进搜索结果、邮件正文这类工具返回内容里,然后拿 Regex 和 Meta 的 Prompt Guard 2 去测。Regex 一条都没抓到,Prompt Guard 2 只抓到 1%。攻击不是直接发给模型的,而是混在模型调用的工具输出里,所以现有检测器基本是瞎的。代码和复现步骤都公开了,但...

Coverage

Follow the reports to see the story from different sides.

Sep 24
  1. Hacker News front pagePick
    Open-source prompt-injection detectors catch 0–1% of realistic AI agent attacks buried in tool output

    This benchmark hides 629 AgentDojo injection attacks inside tool outputs and tests Regex vs. Meta Prompt Guard 2. Regex catches 0%, Prompt Guard 2 catches 1%. The attacks aren't sent directly to the model—they're buried in search results, email bodies, and similar tool responses, so current detectors are effectively blind. Code and reproduction steps are public; the post doesn't include comparisons with commercial detectors.