FrontierCode benchmark sets a new AI coding evaluation bar, with top maintainer approval at 13.4%
FrontierCode 基准测试:AI 编程评估新标准--维护者审核通过率最高仅 13.4%
Cognition released FrontierCode, a coding benchmark built from 150 tasks by more than 20 open-source maintainers and judged against over 3,000 rules, with Claude Opus 4.8 reaching 13.4% approval in the hardest tier and GPT-5.5 reaching 6.3%.
Why it matters: HKR-H/K/R all pass: FrontierCode has a strong 13.4% hook, concrete maintainer-built methodology, and clear coding-agent resonance. Single-source benchmark news keeps it in the 78–84 band, not must-write territory.