After SWE-Bench Pro Saturation, Someone Built a New Benchmark
SWE-Bench Pro 饱和之后,有人做了一把新尺子
DeepSWE says SWE-Bench Pro lost discrimination because of data contamination and verifier flaws; the same model set showed a 62-point spread on the benchmark, while the snippet does not disclose the audited models or test protocol.
Why it matters: HKR-H/K/R all pass: the “new ruler” hook, contamination/verifier claims, and 62-point spread give this real signal for code-agent evaluation. Source reach and impact are below the 85+ same-day tier.