Frontier AI labs still won’t say how they’d contain a rogue model
Frontier AI labs still won’t say how they’d contain a rogue model
Guidelight AI Standards graded five leading labs on their public containment plans for a rogue AI. OpenAI scored highest; Anthropic and Meta came last. Most labs have published almost nothing on what access gets cut or when the system gets shut down if an AI tries to subvert human control. The gap matters as agentic AI takes on more real-world tasks.
Why it matters: A third-party scorecard on rogue-model containment plans turns safety talk into comparable numbers. Anthropic and Meta at the bottom will spark community debate. Score capped below 85 because Guideline isn't a tier-1 evaluator and the article doesn't disclose scoring methodolo...