Skip to content
OpenAI News

OpenAI simulates real-world deployment to catch undesired model behavior before release

Predicting model behavior before release by simulating deployment

OpenAI replays recent real conversations through a candidate model before release, then checks for new undesired behaviors. Across GPT‑5‑Thinking deployments, this Deployment Simulation gave more accurate frequency estimates than traditional evals, surfaced novel misalignment, and reduced the chance models could tell they were being tested. It also works for agentic rollouts with tool use. The method can’t catch behaviors rarer than 1 in 200,000 messages.

Why it matters: OpenAI published a concrete safety-testing method with a paper and reproducible workflow ahead of the GPT-5-Thinking release — not just a vague 'we did safety testing.' The method carries real information gain and hits the concerns of alignment practitioners. Not scored higher...

Read the original ↗Export Markdown