Skip to content
Trending storyDeveloping

Anthropic red team: GLM-5.3 hits full control-flow hijack in 4% of trials, Claude Mythos in 6%

1 report1 sourceupdated 2 hours ago

What happened

AI digest

Anthropic's Frontier Red Team evaluated several models on 100 randomly selected tasks from an internal binary exploitation benchmark. In a report dated 2026-09-30, GLM-5.3 achieved full control-flow hijacking in 4% of trials, and Claude Mythos Preview in 6%. The report says earlier models, including Claude Opus 4.6 and GLM-5.2, failed all of these tasks, and concludes that a meaningful capability threshold has been crossed.

Written by AI from the coverage · updated 53 minutes ago

Coverage

Follow the reports to see the story from different sides.

Sep 30
  1. Simon Willison
    Quoting Anthropic Frontier Red Team

    Anthropic Frontier Red Team 在内部 Binary Exploitation 基准的 100 个随机任务上评测多个模型,GLM-5.3 在 4% 的试验中实现了完整控制流劫持,Claude Mythos Preview 为 6%。报告指出,Claude Opus 4.6 和 GLM-5.2 等更早的模型在这些任务中均未成功,认为一个有意义的能力门槛已被跨过。

Heat over time

Not enough continuous observations to draw a trend yet.