Gray Swan founders: AI security is not just “cybersecurity with AI”
Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan
OpenAI board member Zico Kolter and Gray Swan CEO Matt Fredrikson explain why AI security needs a different mindset. They helped test Anthropic's Mythos model card using their own tool Shade. The core argument: prompt injection creates a new exploit class for computer-use agents, and traditional cybersecurity approaches fall short. Their specialized red-teaming models already beat humans at breaking AI systems. Bigger models don't automatically become more robust. They also cover agent identity, permissions, enterprise guardrails, and AI insurance. The first major prompt-injection breach may be a gray swan—an event everyone can see coming.
Why it matters: Zico Kolter speaking as an OpenAI board member about Anthropic's model safety is a rare, high-authority angle. The core claim—prompt injection as a novel attack surface for agents—is backed by concrete tooling (Shade) and the Mythos model card, not just opinion. The deduction ...