Snorkel releases Senior SWE-Bench, a benchmark that evaluates coding agents like senior engineers
Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
This benchmark gives agents natural-language feature requests and bug reports instead of over-specified specs. A validation agent writes behavioral tests to check solutions, and a taste-scoring system grades whether the code fits the codebase's actual style. The dataset and scoring approach are open-sourced on GitHub.
Why it matters: A new benchmark that directly challenges SWE-bench's setup — natural language tasks instead of over-specified specs, plus a validation agent and taste scoring. Useful for anyone evaluating coding agents. Not an 85 because it just launched with no cross-source buzz or top-model...