Skip to content

Anthropic / Claude

Everything Anthropic: the Claude models, Claude Code, its safety research agenda and company news.

Latest picks

1321–1322 of 1,322

Jan 23, 2025Thursday

OpenAI News

Computer-Using Agent

OpenAI released a research preview of Computer-Using Agent on Jan 23, 2025, and is exposing it first through Operator to U.S. ChatGPT Pro users. The model combines GPT-4o vision with RL-based reasoning and acts through screenshots, a mouse, and a keyboard; it scored 38.1% on OSWorld, 58.1% on WebArena, and 87.0% on WebVoyager. The key point is API-free GUI control, while sensitive actions still require user confirmation.

Why it matters: OpenAI released a research preview of the CUA agent, which uses vision and reinforcement learning to read screenshots and act through a virtual mouse and keyboard, first for U.S. ChatGPT Pro users.

Jan 22, 2025Wednesday

OpenAI News

Trading Inference-Time Compute for Adversarial Robustness

OpenAI reports that o1-preview and o1-mini often drive adversarial attack success rates close to zero as inference-time compute increases. The paper tests math tasks, SimpleQA prompt injection, Attack Bard images, and StrongREJECT misuse prompts; it labels the result as preliminary, and the truncated post does not fully disclose all failure cases. The key point is that this gain comes from longer reasoning at inference, not adversarial training.

Why it matters: OpenAI tested o1-preview and o1-mini on math problems, prompt injection, adversarial images and jailbreak prompts, and attack success rates dropped sharply as inference-time compute rose.