OpenAI moved Codex onto the desktop and aimed it at more than 3 million weekly developers. My read is blunt: this is not a feature refresh. It is OpenAI turning a code assistant into a persistent software agent, with the developer workstation as the beachhead.
The article gives enough hard facts to take that seriously. Codex can now use a Mac with background computer use by seeing the screen, clicking, and typing with its own cursor. Multiple agents can run in parallel without blocking the user’s own apps. The app now adds an in-app browser, image generation through gpt-image-1.5, SSH access to remote devboxes in alpha, multiple terminal tabs, PR review workflows, rich previews for files like PDFs and spreadsheets, automations that can reuse threads, scheduled future work, and a preview of memory. OpenAI says it is also shipping 90-plus new plugins spanning skills, app integrations, and MCP servers, including Atlassian Rovo, CircleCI, CodeRabbit, GitLab Issues, Microsoft Suite, Neon by Databricks, Render, and others. Availability starts now for signed-in ChatGPT users on the Codex desktop app. Personalization and memory are still rolling out unevenly, and the body does not disclose pricing, permission granularity, latency, success rates, or sandbox boundaries.
That missing data matters because the strategic move here is bigger than the product copy suggests. OpenAI is making a play for the work layer around code, not just code generation itself. The old Copilot frame was “help me write functions faster.” This Codex frame is “stay with me across the full software lifecycle, touch the tools that do not expose APIs, remember how I work, and keep going after I walk away.” That is a different product category. It pushes Codex closer to the same broad lane as Devin, Cursor’s agent modes, Windsurf’s workflow pitch, and Anthropic’s computer-use demos. The difference is distribution. OpenAI already has a huge ChatGPT surface and is now trying to convert that reach into an agent shell for developers.
I have thought for a while that developers would be the first group where desktop agents actually stick, not because they are easier to sell to, but because their work is unusually verifiable. A PR merges or it does not. CI turns green or it does not. A frontend matches the design or it does not. That feedback loop is gold for agent products. If Codex can open the browser, run tests, inspect screenshots, address review comments, and continue tomorrow with memory intact, then OpenAI is no longer competing on autocomplete quality alone. It is competing on who owns the execution loop.
This is also where I start pushing back on the marketing. “Codex for almost everything” is too broad for what is actually disclosed. The concrete use cases in the body are still developer-heavy: frontend iteration, games, PR comments, SSH, JIRA, Slack, Gmail, Notion. Yes, the product is moving beyond coding. No, this does not yet read like a general-purpose computer agent in the strong sense. The hard problems are permissions, safety, and recovery, and the post is thin there. Codex can see the screen, click, type, wake itself up later, and remember context across time. Fine. Then how does it handle 2FA prompts, secrets, production commands, internal admin tools, and personal documents? Are permissions app-level, task-level, workspace-level, or account-wide? Is memory local, organizational, or cloud-shared by default? The article does not say. For enterprise buyers, that omission is not cosmetic. It is the whole decision.
I also have some doubts about the computer-use claim until OpenAI shows reliability numbers. Anthropic’s early computer-use push made the category legible, but the gap between a strong demo and a stable daily workflow was obvious. Speed, misclicks, window focus, brittle selectors, and failure recovery all bite fast. OpenAI emphasizes that multiple agents can work in parallel on your Mac without interfering with your own work. Good idea. But the body gives no task completion rate, no median duration, no rollback behavior, and no explanation of how parallel agents avoid stepping on each other at the OS level. Without those, the feature reads like a promising shell around an unproven operations model.
The broader pattern is still important. Model vendors are drifting upward into applications and downward into operating surfaces at the same time. OpenAI is not just fighting GitHub Copilot for completion share anymore. It is reaching into the browser layer, the local desktop layer, the ticketing layer, and the remote environment layer. Once a vendor wins that workflow graph, the underlying model becomes somewhat less sticky than the surrounding execution environment. Teams can swap models more easily than they can rip out an agent that already knows their repos, tickets, docs, terminals, and approval paths.
That is why I would not judge this release by asking whether Codex got better at coding benchmarks. OpenAI did not even center benchmarks here, which is telling. I would judge it on two operational questions the article leaves open. First, can an admin disable memory, browser control, SSH, and cross-app access with real granularity. Second, what auditing and rollback exist when a scheduled task keeps running across days. If OpenAI has strong answers there, Codex starts to look like infrastructure for software work. If not, it stays a very polished personal demo with a lot of ambition wrapped around it.