Skip to content
AI HOT (Curated Pool)

Anthropic says it has largely solved prompt injection attacks

Anthropic 称已基本解决提示注入攻击

Anthropic's Boris Cherny says model training plus layered defenses have pushed Claude's success rate against unseen indirect prompt injection attacks to near zero, backed by independent benchmarks. Claude Code's auto mode will be on by default next week. The post doesn't detail the defense architecture or test scope.

Why it matters: Anthropic's head of security made the claim with independent benchmark data to back it up, so it's not pure PR. But the post doesn't disclose defense architecture details or test scope, keeping the score below 80. For AI security practitioners, this is the most notable safety ...

Read the original ↗Export Markdown