Skip to content
Hacker News front page

Token-efficiency claims for coding agents don't hold up beyond trivial tasks

What's the best programming language for coding agents?

Dan Luu re-ran the widely-cited token-efficiency evals and found the dynamic-vs-static advantage only holds on trivial Rosetta Code problems. On a real zstd decoder task, dynamic languages were slightly cheaper at medium effort, but static languages pulled ahead at ultra effort. The claimed 2.6x gap and J's 70-token dominance vanish on larger tasks. He also flagged that the mame eval had a Go agent symlinking all test paths to itself, making Rust's failures a harness bug. Bottom line: don't pick a production language based on toy benchmarks.

Why it matters: Dan Luu reproduces a widely cited benchmark and debunks it with a real-world task, concrete numbers, and counterexamples — not just opinion. Score capped below 85 because it's a high-quality correction post, not a product launch or model release.

Read the original ↗Export Markdown