Skip to content
Trending storyPast story

METR's Predeployment Evaluation of Claude Opus 5.5

1 report1 sourceupdated 6 days ago

What happened

Summary

METR 在模型发布前测了 Claude Opus 5.5 对 AI 研发的影响。结论是它比 Fable 5.1 强一点,在预算版 NanoGPT 速通赛、游戏机器人这类可验证任务,以及 LMCA 概念推理、Sunlight 开放研究这类难验证任务上都有提升,但不是断崖式飞跃。Anthropic 内部问卷也说它延续了 Mythos 级别的趋势。模型在需...

Coverage

Follow the reports to see the story from different sides.

Sep 22
  1. AI HOT (Curated Pool)Pick
    METR's Predeployment Evaluation of Claude Opus 5.5

    METR evaluated Claude Opus 5.5's impact on AI R&D. It's a modest step up from Fable 5.1, not a leap toward full automation. Gains showed on verifiable tasks like Budget NanoGPT and Gaming Bot, and on harder-to-verify ones like LMCA and Sunlight. Anthropic's internal questionnaire says it continues the Mythos-level trend. A separate, undisclosed METR report estimates AI already accelerated Anthropic's overall R&D by ~1.5x, with a 30% chance of 2x. The post doesn't disclose specific parameters, pricing, or a release timeline.