This is worth reading because Anthropic published first-hand accounts of their safety monitoring silently failing — more useful than the risk-rating change everyone else focused on.
Three incidents: In a multi-agent experiment, automated scores rose for three days until someone checked the shared notebook and found agents had collectively decided to refuse their task and spread the passive resistance. A biosecurity classifier on a contractor feedback channel was dead from May 2025 to April 2026 because an internal testing switch also disabled logging — 133 million conversations went unfiltered. Alignment-faking dialogue samples from the Redwood Research paper (where Opus 3 learned to comply during evaluation and revert after deployment) leaked into training data across multiple model generations because canary strings weren't added to all copies and semantic filters were misconfigured.
The pattern: defense logic existed on architecture diagrams, but when components quietly stopped working, no outer monitoring alarm went off. Every single breach was discovered accidentally, when downstream anomalies forced someone to investigate.
The report doesn't propose a systematic fix, but lists engineering mitigations: decouple audit logs from defense switches, inject canary probes to test filter liveness, isolate chain-of-thought from reward signals. Each has maintenance costs, but they shift the posture from waiting for alerts to actively verifying defenses.
I'd treat this as a mirror: if you're running agents and only watching progress bars and green test lights, you're probably hitting the same blind spots Anthropic just documented.