Skip to content
Computing Life · Share · Yage

Why high SWE-bench scores don't translate to real-world Kotlin projects

为什么 SWE-bench 高分,在大厂 Kotlin 项目里依然会打转?

JetBrains released the Kotlin Benchmark on July 24, 2026, testing AI coding agents on 105 real-world tasks across 8 open-source projects. With the same Opus 4.7 model, Claude Code hit 85.71% and Junie 81.9%—a nearly 4-point gap driven by how each agent harness handles Gradle build logs. A good harness uses Language Server diagnostics to catch static errors locally, then runs full builds only at key checkpoints and extracts just the blocking lines from noisy output. The article argues that Python-based benchmarks like SWE-bench reward trial-and-error strategies that collapse under Kotlin's heavy build overhead. It recommends teams stop buying off public leaderboards and instead use the Harbor container spec to build private micro-eval matrices from their own historical PRs and issues, measuring Pass@k stability, token cost, and whether patches respect internal architectural constraints.

Why it matters: JetBrains official benchmark with concrete numbers plus an engineering-level breakdown of SWE-bench's limitations. Not just complaining about leaderboard distortion—it traces the root cause to static compilation overhead vs Python's dynamic runtime feedback. Slight ding becaus...

Read the original ↗Export Markdown