Skip to content
AI HOT (Curated Pool)

OpenAI releases GPT-Red: an automated red-teamer that makes GPT-5.6 far more resistant to prompt injection

OpenAI 发布 GPT-Red:通过自动化红队测试提升模型鲁棒性

OpenAI trained GPT-Red, an automated red-teaming model that finds vulnerabilities through self-play and feeds the attacks into adversarial training for GPT-5.6. The result: GPT-5.6 Sol shows 6x fewer failures on the hardest direct prompt injection benchmark compared to the best production model from four months ago. OpenAI says this is the first time they've dedicated compute at the scale of their largest post-training runs purely for safety. The post includes case studies—exfiltrating internal files, forwarding API keys, running malicious build scripts—but does not disclose GPT-Red's parameter count or detailed training recipe.

Why it matters: OpenAI's first safety run at main-model training scale, with a concrete 6× failure reduction on direct prompt injection for GPT-5.6 Sol. A lab-grade safety result that red-teaming and deployment teams will benchmark against. Not scored higher because it's a single-source blog ...

Read the original ↗Export Markdown