Skip to content
AI HOT (Curated Pool)

METR's Predeployment Evaluation of Claude Opus 5.5

METR 发布 Claude Opus 5.5 部署前评估摘要

METR evaluated Claude Opus 5.5's impact on AI R&D. It's a modest step up from Fable 5.1, not a leap toward full automation. Gains showed on verifiable tasks like Budget NanoGPT and Gaming Bot, and on harder-to-verify ones like LMCA and Sunlight. Anthropic's internal questionnaire says it continues the Mythos-level trend. A separate, undisclosed METR report estimates AI already accelerated Anthropic's overall R&D by ~1.5x, with a 30% chance of 2x. The post doesn't disclose specific parameters, pricing, or a release timeline.

Why it matters: METR's pre-deployment eval of Claude Opus 5.5 brings an independent third-party lens with concrete task comparisons. Not scored higher because the finding is 'incremental, not a leap,' limiting impact, but as a safety/capability crossover assessment for an Anthropic model, it'...

Read the original ↗Export Markdown