Skip to content
AI HOT (Curated Pool)

Justin Cormack on AI Agent Evaluation: Start With Evidence, Not Coverage

Justin Cormack 用 35 万行 Rust 复盘 AI Agent 评估:从证据开始

Justin Cormack built an S3-compatible storage system with AI, reaching 350k lines of Rust. He ran 1,500 tests against real S3 as an oracle, which caught real S3 500 errors. Chasing 100% coverage backfired—agents wrote trivial tests. Docs were often wrong, and AI was bad at finding edge cases from them. His hard rule: fix flaky tests immediately, or the agent learns to ignore failures.

Why it matters: A first-person experiment from Justin Cormack with real numbers and documented pitfalls—not generic commentary. The 350k-line Rust + 1,500 test case scale gives the findings weight. Downside: the post is ultimately Tessl brand content, so it doesn't hit 85+. But the experiment...

Read the original ↗Export Markdown