BackendForge turns AI coding benchmarks from written exams into interviews, halving pass rates on the same code
BackendForge:AI 已经能写完整 App,评测也得学会追问
BackendForge starts with 7,250 tests across 56 backend tasks, then has agents hunt for gaps and add 640 more. Only +8.8% more tests, yet GPT-5.5's passing tasks drop from 31 to 16, and Claude Opus 4.7 from 33 to 10. Same code, sharper questions on permissions, dirty data, and cascading effects. The paper says materials will be released, but they aren't public yet—don't generalize these numbers yet.
Why it matters: BackendForge raises a sharp benchmarking methodology question: when agents can write full backends, tests must learn to probe permissions, dirty data, and operation ordering. The numbers are hard—GPT-5.5 and Claude Opus 4.7 saw their full-pass rates halved under the new test s...