GLM-5.1 beat Opus 4.6, GPT-5.4, and Gemini 3.1 Pro on SWE-Bench Pro
GLM-5.1 scored 58.4 on SWE-Bench Pro, ahead of Opus 4.6 at 57.3, GPT-5.4 at 57.7, and Gemini 3.1 Pro at 54.2. The post also says it is an MIT-licensed open-weight model; the post does not disclose eval setup, cost, or whether all models were tested under identical conditions. Watch reproducibility, not a single leaderboard snapshot.
Why it matters: Open-weight GLM-5.1 beating closed leaders on SWE-Bench Pro is a real hook, and the score deltas are concrete. Source authority is weak: this is a single X post with no disclosed eval setup, cost, or equal-condition proof, so it stays low-featured rather than higher.