Skip to content
Hacker News front page

Snorkel releases Senior SWE-Bench, a benchmark that evaluates coding agents like senior engineers

Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

This benchmark gives agents natural-language feature requests and bug reports instead of over-specified specs. A validation agent writes behavioral tests to check solutions, and a taste-scoring system grades whether the code fits the codebase's actual style. The dataset and scoring approach are open-sourced on GitHub.

Why it matters: A new benchmark that directly challenges SWE-bench's setup — natural language tasks instead of over-specified specs, plus a validation agent and taste scoring. Useful for anyone evaluating coding agents. Not an 85 because it just launched with no cross-source buzz or top-model...

Read the original ↗Export Markdown