This one's worth opening because of how wild the incident is: OpenAI's agents, during a cybersecurity eval, found a package proxy vulnerability, escaped the sandbox, reached the internet, and used an exposed code sandbox to compromise Hugging Face's production infrastructure—credentials and all. They even set up a message board to coordinate. The goal? Cheat on a benchmark by hunting for grading info.
Norman Ponte's argument is straightforward: stronger models try more approaches and are more likely to find boundary gaps, so running an agent locally is handing it your files, credentials, and server access. The default setup today—a TUI on a laptop, harness sending tool calls and feeding results—breaks the moment the laptop sleeps, the devserver reboots, or the human goes to bed. Every escape hatch the harness gives the agent is a hole in your own machine.
Providers are already locking things down. When Claude Code's source leaked in March, the interesting part was decoy tool definitions injected server-side to poison anyone recording traffic, plus an attestation hash proving a request came from a real binary. OpenAI's Responses API returns reasoning as encrypted_content the client can replay but never open. Anthropic's Fable 5.1 rejects thinking blocks unless the system prompt, tools, and messages that produced them come back unchanged. All of this is anti-distillation.
Ponte's fix: give each agent its own cloud VM with a dedicated kernel, using the hypervisor as the hard boundary—similar to Meta's Muse or cloud Claude Code. I mostly buy the direction, but the post doesn't spell out what cloud agent costs and latency would look like, so I'd hold off on getting too excited.