Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,488 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

1481–1488 of 1,488

Feb 3, 2025Monday

OpenAI News

Introducing deep research

OpenAI launched deep research in ChatGPT, an agentic feature that spends 5 to 30 minutes finding, analyzing, and synthesizing hundreds of web pages, images, and PDFs into a cited report. It runs on a version of OpenAI o3 optimized for web browsing and data analysis and was trained on real-world browser and Python tasks; after the April 2025 update, Plus/Team/Enterprise/Edu get 25 queries per month, Pro 250, and Free 5. The key point is a productized workflow for multi-step, source-backed research, not a basic search refresh.

Why it matters: OpenAI launched deep research in ChatGPT, where an o3 variant tuned for browsing and data analysis reads hundreds of pages and returns a cited report in 5 to 30 minutes.

Jan 23, 2025Thursday

OpenAI News

Operator System Card

OpenAI published the Operator System Card on Jan 23, 2025 and said its Computer-Using Agent can be deployed only if its post-mitigation score is Medium or lower. The card rates CBRN, cybersecurity, and model autonomy as Low, and persuasion as Medium; it highlights harmful tasks, model mistakes, and prompt injection. The key mechanism is human confirmation plus task refusal: critical steps like financial transactions, emails, and calendar deletion need approval, while stock trading is fully restricted.

Why it matters: OpenAI published the Operator system card, capping deployment at post-mitigation Medium and requiring human confirmation for high-risk actions like financial transactions, sending email and deleting calendar events, while banning stock trades outright.

OpenAI News

Computer-Using Agent

OpenAI released a research preview of Computer-Using Agent on Jan 23, 2025, and is exposing it first through Operator to U.S. ChatGPT Pro users. The model combines GPT-4o vision with RL-based reasoning and acts through screenshots, a mouse, and a keyboard; it scored 38.1% on OSWorld, 58.1% on WebArena, and 87.0% on WebVoyager. The key point is API-free GUI control, while sensitive actions still require user confirmation.

Why it matters: OpenAI released a research preview of the CUA agent, which uses vision and reinforcement learning to read screenshots and act through a virtual mouse and keyboard, first for U.S. ChatGPT Pro users.

OpenAI News

Introducing Operator

OpenAI released Operator on Jan 23, 2025 as a research preview for U.S. Pro users; it uses its own browser to click, type, and scroll through web tasks. It runs on Computer-Using Agent, combining GPT-4o vision with RL-based reasoning; the post says it sets SOTA on WebArena and WebVoyager but does not disclose scores. The key boundary is control: login, payment, and CAPTCHA flows hand control back to users, and a July 17 update says it was folded into ChatGPT agent.

Why it matters: OpenAI released a research preview of Operator, whose CUA model drives browser buttons and menus directly without site APIs, handing control back to the user at logins, payments and CAPTCHAs.

Oct 10, 2024Thursday

OpenAI News

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

OpenAI released MLE-bench, a benchmark built from 75 Kaggle competitions to measure ML engineering ability in AI agents. The best setup, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level on 16.9% of tasks; the benchmark code is open-source.

Why it matters: OpenAI builds MLE-bench from 75 real Kaggle competitions so agents run the full pipeline, with the best setup reaching bronze-medal level in 16.9% of contests.

Oct 1, 2024Tuesday

OpenAI News

Introducing the Realtime API

OpenAI launched a public beta of the Realtime API on Oct. 1, 2024 for all paid developers, using a persistent WebSocket to stream low-latency speech-to-speech interactions with GPT-4o. It supports function calling and interruption handling, priced at $5/1M text input tokens and $100/1M audio input tokens; the post also says audio I/O for Chat Completions would arrive in the following weeks.

Why it matters: Removes the speech-to-text and text-to-speech hops that made voice agents slow.

Aug 6, 2024Tuesday

OpenAI News

Introducing Structured Outputs in the API

OpenAI released Structured Outputs on Aug 6, 2024, making model outputs conform to developer-supplied JSON Schemas; `gpt-4o-2024-08-06` scored 100% on complex schema-following evals versus under 40% for `gpt-4-0613`. The feature is enabled with `strict: true` in function calling and works on tool-supporting models including `gpt-4-0613`, `gpt-3.5-turbo-0613`, and later. The key shift is constrained decoding plus schema training, not just valid JSON from JSON mode.

Why it matters: Structured Outputs makes models follow a developer's JSON Schema exactly, and the accuracy gap between new and old models shows why schema adherence is a model capability, not a prompt trick.

Feb 13, 2024Tuesday

OpenAI News

Memory and new controls for ChatGPT

OpenAI says ChatGPT is getting memory and new controls, with 2 changes disclosed in the title. The body is empty, so default state, opt-out scope, and user-tier availability are not disclosed. The key issue is control granularity; the title alone is not enough to judge product impact.

Why it matters: OpenAI gave ChatGPT memory across chats, so it keeps user preferences and past details, with controls to turn memory off, tell it to forget, and use temporary chats.