Trading Inference-Time Compute for Adversarial Robustness
OpenAI reports that o1-preview and o1-mini often drive adversarial attack success rates close to zero as inference-time compute increases. The paper tests math tasks, SimpleQA prompt injection, Attack Bard images, and StrongREJECT misuse prompts; it labels the result as preliminary, and the truncated post does not fully disclose all failure cases. The key point is that this gain comes from longer reasoning at inference, not adversarial training.
Why it matters: OpenAI tested o1-preview and o1-mini on math problems, prompt injection, adversarial images and jailbreak prompts, and attack success rates dropped sharply as inference-time compute rose.