Skip to content
Hacker News front page

SWE-rebench leaderboard: 13 models and 4 agents benchmarked on real-world bug fixes across Go, Java, Python, Rust, and TypeScript

13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS

Nebius built this benchmark using 111 real GitHub issues from 65 repos across Go, Java, Python, Rust, and TypeScript. Fable 5 leads with a 64.5% resolved rate at $4.40 per problem. Grok 4.5 and Opus 5 both hit above 63%, but Grok 4.5 costs only $1.47 per problem—much cheaper. Among agents, Junie scores 61.8% at $0.81, while Claude Code gets 60.4% at $3.39. DeepSeek-V4 Pro resolves 40.2% at just $0.15 per problem, the cheapest in the top 14. The post does not break down per-language performance or explain why many models—from Claude Opus 4.1 through Sonnet 4.6—are listed as N/A.

Why it matters: Nebius built this benchmark from 111 real GitHub issues across 65 repos, mixing models and agents with transparent cost data. All three HKR axes hit, but it's a third-party eval, not a model release—caps below 85.

Read the original ↗Export Markdown