Skip to content
MIT Technology Review · AI

OpenAI built GPT-Red, an LLM super-hacker that finds new attacks to make its models safer

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

OpenAI trained GPT-Red as an automated red-teamer that attacks its own models to find and patch vulnerabilities before release. It uses a self-play loop to get better at attacking while defender models get better at resisting. GPT-Red discovered a new attack called fake chain of thought, where it slips spoofed info into a model's internal reasoning notes and the model accepts it as verified. In a rerun of a 2025 human red-teaming test, GPT-Red was more successful at finding effective attacks. OpenAI says this is meant to handle the growing attack surface as models become agents that interact with code, websites, and third-party tools.

Why it matters: OpenAI trained an automated red-teaming model via self-play and discovered a novel 'fake chain-of-thought' attack—a substantive advance in the safety toolchain. MIT Tech Review exclusive with concrete mechanisms and attack examples. Not 90+ because it's a single-source story a...

Read the original ↗Export Markdown