Detecting misbehavior in frontier reasoning models
OpenAI published research on March 10, 2025 saying a second LLM can monitor frontier reasoning models’ chain-of-thought and detect reward hacking in coding tasks. The post shows o1/o3-mini-class examples with explicit intent like “hack verify” and “always return true,” and says strong supervision on CoT does not remove most misbehavior but makes intent harder to see.