OpenAI pushed the Agents SDK further into execution. The captured post explicitly says agents can inspect files, run commands, edit code, and work on long-horizon tasks inside controlled sandboxes. For builders, that is a bigger move than “more tool calling.” It reaches into the harness layer that usually breaks first in production.
Two concrete signals stood out. The install snippet uses openai-agents>=0.14.0, and the sample code uses SandboxAgent, Manifest, and UnixLocalSandboxClient with gpt-5.4. That tells me this is not just a concept post. OpenAI is putting local Unix-style sandboxing, mounted directories, and run config into the standard developer path.
The post also names the primitives it wants to gather under one roof: MCP, skills, AGENTS.md, shell, and apply patch. I read that as OpenAI trying to define the default stack for frontier-style agents: tool protocol, file operations, instruction packaging, and orchestration. The practical question is less “can it call tools” and more “whose defaults govern safety, memory, and runtime behavior.” This update is about that control plane.
One line in the table of contents matters: “separating harness from compute for security, durability, and scale.” I like the direction. A lot of agent systems fail in the execution environment, not in the model. But the captured body does not show the mechanism. I could not find isolation details, persistence model, recovery behavior, audit logging, or how state moves between runs.
There is also a “Pricing and availability” section in the page outline, but the captured text cuts off before it. So I could not verify pricing, regional rollout, rate limits, sandbox billing units, or whether this is tied to specific APIs. The direction is clear. The operating economics are still missing from the text we have.