OpenAI made the top-line claim first: o3 and o4-mini stayed below the High threshold in all 3 tracked risk categories—bio/chemical, cybersecurity, and AI self-improvement. The problem is that this page is basically a wrapper, not the system card itself. It says Preparedness Framework V2 was used and the Safety Advisory Group reviewed the results, but it does not expose the parts practitioners actually need: sub-scores, evaluation setups, pass/fail criteria, refusal boundaries, tool gating conditions, or how browsing and Python changed the threat model.
My read is pretty simple: this is not a meaningful increase in safety transparency. It is a conclusion published ahead of the evidence. For a general audience, “below High” sounds reassuring. For anyone who has had to evaluate model risk, it is thin. That matters more here because OpenAI is not talking about a plain chat model. It explicitly says o3 and o4-mini have full tools: web browsing, Python, image and file analysis, image generation, automations, file search, memory. Risk posture for a tool-using model is different from risk posture for a text-only model. Once the model can browse, run code, transform images, and chain steps together, evaluation has to move from “can it answer a dangerous question” to “can it acquire, verify, and execute operationally useful steps.”
There is also a broader pattern here. Over the last year, system cards across the frontier labs have become less controversial at the headline level and more controversial at the methods level. The fight is no longer over whether labs publish a safety document. The fight is over whether outside researchers can inspect the benchmark design, threat assumptions, and mitigation details well enough to disagree. From memory, Anthropic’s stronger system card releases usually gave more behavioral examples and more visible mitigation framing, even when the company still kept important internals private. This OpenAI page, by contrast, is a sparse announcement page that pushes the substantive details into a linked document. That does not prove the underlying card is weak. It does mean this page itself should not be treated as serious evidence.
I also have a specific pushback on the framing around deliberative alignment. OpenAI highlights that these reasoning models can reason about safety policies in context. Fine. That approach has merit. But it also depends on the reasoning process and tool-use loop holding up under adversarial conditions. Browsing introduces prompt injection. File analysis introduces hostile payloads and hidden instructions. Multi-step agent behavior introduces error accumulation and policy drift across turns. “Below High” is only meaningful if we know the threat model: single-turn prompts, long-horizon agent tasks, external websites, user-supplied files, or some mix. This page does not say.
So I would read this less as a full safety disclosure and more as OpenAI laying policy and reputational groundwork for broader agent deployment. The positive signal is that tool use is now clearly inside the preparedness frame, and OpenAI is attaching addenda for Codex and Operator-style deployments that sit closer to execution. The weak part is that the public-facing disclosure here is still operating at verdict level. If OpenAI wants technical readers to trust the claim, the linked system card needs to carry the real load: concrete tasks, thresholds, success rates, intervention points, and whether risky tools are constrained by default. On the evidence shown on this page alone, the claim is directionally useful but not yet audit-grade.