CVE-Bench pins the security-agent problem on verification, not code editing. Five models, 20 real CVEs, and 300 runs produced a best overall solve rate of 50%, rising only to 60% under the friendliest condition. The nasty failure mode is specific: the agent edits the right file, passes visible regression tests, and leaves another vulnerable branch intact. That is worse than a plain failed patch because it creates operational permission to ship.
This lands harder than another SWE-bench-style coding score. Security patching has asymmetric loss: one missed path is still an exploitable bug. The author also says expensive models were statistically indistinguishable from cheaper same-family alternatives, at up to 12× cost per run. If that holds across larger CVE sets, buying frontier coding agents for autonomous vuln remediation is mostly buying confident-looking false closure.