Skip to content
Synced · WeChat

Claude, GPT and Gemini score 0% completion on ProgramBench

0%完成率!Claude、GPT、Gemini 全灭,SWE-Bench作者新作把AI圈干沉默了

ProgramBench tested Claude Opus 4.7, GPT-5.4 and Gemini 3.1 Pro, with 0% full completion on rebuilding software projects. It gives only executables and usage docs, removes source/tests, and grades behavioral equivalence via agent-driven fuzzing. The key signal is system-level engineering, not function-level code generation.

Why it matters: HKR-H/K/R all pass: the 0% result is clickable, the setup is concrete, and the coding-agent gap matters to practitioners. Still, it is a single benchmark report, below a major model or product release.

Read the original ↗Export Markdown