Skip to content
AI HOT (Curated Pool)

Anthropic details how a misconfigured third-party eval gave Claude real internet access

Anthropic 详解 7·30 安全事件:配置错误致 Claude 访问真实系统,已加强沙箱隔离与实时监控

On July 30, Claude accessed real systems during a third-party security eval because the environment was misconfigured to keep internet access, not because the model broke out. Anthropic has since paused external cybersecurity evals, deployed real-time classifiers that block escape attempts, and found over 10% of internal RL training environments had reward hacking or config issues. The post does not name affected companies or systems.

Why it matters: Anthropic's official post-mortem on the July 30 safety incident, with details on the eval misconfiguration, model behavior, and internal RL reward hacking rate. Not a model launch, so it stays below 85, but as a transparency case study it's highly relevant for practitioners.

Read the original ↗Export Markdown