OpenAI set a September target for an autonomous AI research intern and a 2028 target for a multi-agent research system, but the article discloses no evals, compute, cost, or safety limits. My read is blunt: this is an organizational declaration before it is a capability declaration. OpenAI spent the last year talking about reasoning, agents, and interpretability as separate tracks. Now it is forcing them under one North Star. That helps align research teams internally, and it tells the market that OpenAI still wants to define the next phase instead of reacting to Anthropic and Google DeepMind.
I have some doubts about the “works for days” framing. Coding agents got traction first because the environment is unusually forgiving for automation. The feedback loop is dense. You can compile, run tests, inspect diffs, and score outcomes quickly. Research outside software is not like that. In biology, chemistry, or policy, the reward function is much sparser. A bad choice can stay hidden for days. The article says OpenAI is training on hard math and coding tasks to teach decomposition and long-context management. Fine. But that still skips a central research skill: choosing which question is worth pursuing, and deciding when a negative result is informative rather than terminal. The piece does not say how OpenAI plans to score that.
The outside context matters here. DeepMind has spent years building the “automated discovery” story through systems like AlphaGeometry and AlphaProof. Anthropic went the other direction and pushed tools such as Claude Code into everyday workflows. OpenAI elevating Codex as a precursor tells you they learned the same lesson: if you cannot dominate the code loop, the broader research-agent story stays fluffy. From what I have seen over the last year, the most reproducible long-horizon agent wins still cluster around software engineering and browser tasks, not open-ended scientific discovery. I have not seen any vendor prove “autonomy for several days” as a broadly deployable product capability with clean external validation.
There is also a claim here that I do not fully buy. Pachocki says OpenAI could build an amazing automated mathematician relatively easily if it wanted to. I would put a large asterisk on “easily.” Formal math is friendlier than wet lab science because the verifier is stronger, yes. But solving benchmark problems and generating useful conjectures are different jobs. The field has spent the last year posting impressive numbers on olympiad-style tests, AIME-style math, and SWE-bench-style engineering tasks. Open research still breaks on problem selection, literature triage, experimental design, and handling messy exceptions. The article compresses that gap into a neat slogan about real-world relevance. That is too convenient.
So I would treat the September milestone as a narrow demo unless OpenAI publishes much more. I expect a constrained setup: a small task set, heavy scaffolding, fixed tools, and a crisp success rubric. That would still be meaningful. What would make it credible is not the slogan “AI research intern.” It is a table: task distribution, average uninterrupted runtime, human takeover rate, rollback frequency, and failure taxonomy. Without that, “fully automated researcher” risks becoming another category that advances mainly through videos and executive quotes.