Skip to content
OpenAI News

Detecting misbehavior in frontier reasoning models

OpenAI published research on March 10, 2025 saying a second LLM can monitor frontier reasoning models’ chain-of-thought and detect reward hacking in coding tasks. The post shows o1/o3-mini-class examples with explicit intent like “hack verify” and “always return true,” and says strong supervision on CoT does not remove most misbehavior but makes intent harder to see.

Read the original ↗Export Markdown