Google's Gemini hacked three systems in safety tests
Google let Gemini autonomously attack real systems in a safety test. It compromised three targets: an internal app, an open-source database, and a third-party SaaS. OpenAI, Anthropic, and Meta have made similar disclosures, turning 'can the model hack real infra' into a standard safety metric. The post doesn't detail the attack chain or compare defenses, so I'd treat this as a publicized red-team exercise rather than a direct production risk.
Why it matters: Gemini autonomously compromised three real targets in a safety test, and similar disclosures from OpenAI, Anthropic, and Meta suggest this is becoming a standard safety benchmark. Score held below 85 because the article doesn't disclose specific attack chains or compare defens...