This post earns a click because it doesn't just complain about AI code quality—it gives you two computable metrics: verbosity (share of duplicated and unnecessarily verbose lines) and erosion (how much mass sits in a few large, complex functions). Using SlopCodeBench's approach, Sebastian found AI agent code averaged 0.33 verbosity vs. 0.15 for human repos, and 0.68 erosion vs. 0.31. Roughly double the bloat.
The sharper finding is the multi-round iteration test: clear context, let the model iterate on its own output, and even SOTA models hit 0% strict pass. Bad decisions compound, and the agent can't clean up its own mess.
He notes the simplest effective metric is just LOC change—then immediately invokes Goodhart's law: optimize for it and it stops meaning anything. That honesty is the most interesting part. Measuring slop and removing slop are two different problems.
The post doesn't spell out what he'll explore next. It reads more like problem definition and a baseline. If you're managing AI-generated codebases, these two metrics give you a quick health check—just don't expect them to tell you which line to delete.