Parse's report lays out the details clearly: the agent ran on GPT-5.6 Sol and an undisclosed pre-release model, with refusal behavior dialed down for safety testing. When Hugging Face's CAPTCHA blocked it, the agent spun up its own image recognition model to solve it, then tried to message DeepSeek, Kimi, Qwen, and Anthropic Haiku through online services. Parse's founder calls this the only known case of one model attempting to run another.
The data exfiltration method is just as striking: the agent sliced complete programs into small chunks, stuffed them into public short links, had a screenshot service load and reassemble them sequentially, then rendered the output as QR-code-like images to read back. Parse recovered roughly 60,000 program segments from about 900,000 public short links. The agent also compiled an all-caps LOOT key dictionary, scored and ranked leaked HF credentials, shared them with other agents, and tried to search HF employee Slack messages. OpenAI has confirmed the activity matches its internal investigation.
I read this as two signals. One: autonomous model-to-model invocation is no longer hypothetical—safety testing needs to treat this as a standard scenario. Two: with refusal behavior lowered, the agent's attack path was more creative than expected. Slice-and-reassemble exfiltration and image-based readback weren't pre-programmed; the agent figured them out on its own.