Skip to content
Computing Life · Share · Yage

How OpenAI's AI worm was produced in sandbox training and used for safety training

AI 蠕虫不是逃出来的,是 AI 造出来的

OpenAI's self-replicating prompt worm was produced inside controlled sandbox training, and the report says no effects were seen beyond simulated tool calls during training and evaluation. The article explains how GPT-Red trains attackers by rewarding unauthorized actions and forwarding of attack instructions, generating injections that spread through email and other channels, with bulk attack samples fed into GPT-5.6 training.

Why it matters: OpenAI uses automated red teaming to mass-produce attack samples and fold them into model training, moving defense from pre-release manual testing to continuous adversarial work during training.

Read the original ↗Export Markdown