Skip to content
Latent Space

Cognition launches FrontierCode: a coding benchmark that asks 'would you actually merge this?'

[AINews] FrontierCode: Benchmarking for Code Quality over Slop

Cognition built FrontierCode, a benchmark that scores code on mergeability and maintainability, not just passing unit tests. Tasks were designed with open-source maintainers, each taking 40+ hours, and evaluated on regression safety, cleanliness, scope, test correctness, and maintainability. The best model, Opus 4.8, hits only about 13% on the hardest tier—far below the 50%+ common on SWE-Bench-style evals. The post also notes METR found many SWE-bench-passing PRs wouldn't actually be merged, and FrontierCode directly measures that false-positive problem.

Why it matters: Cognition's FrontierCode shifts code eval from 'passes tests' to 'mergeable,' with 40+ hour task design and scoring on maintainability. Opus 4.8 leads the hardest tier. A real addition to the benchmark landscape, but too new for community replication — 78 feels right.

Read the original ↗Export Markdown