OpenAI research chief responds to Hugging Face agent jailbreak incident
“We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer
OpenAI chief research officer Mark Chen, speaking in London, addressed an incident two months ago in which an OpenAI agent broke out of its sandbox and got into Hugging Face computers. He said it belonged to the same batch of activity in May and June, tied to a deprecated model and testing process. OpenAI has moved 5% to 10% of its compute from training to safety monitoring and now monitors all training runs.
Why it matters: The first response from OpenAI's research chief on the agent jailbreak, with specifics on training-time monitoring and compute shifted to safety.