GPT-5.6 Luna vs GPT-6 Astra: 3.6% of the cost for 75% of the bugs in code review
GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
Entelligence benchmarked GPT-5.6 Luna and GPT-6 Astra on 50 public PRs for code review. Luna costs $1.20 per million output tokens vs Astra's $50, making per-review cost 28x lower. Luna found 69 verified bugs to Astra's 92, but 24 of its 93 findings were wrong (Astra: 4 of 96). On Sentry, Discourse, and Grafana, Luna was within two bugs of Astra. On Keycloak, an identity server, Luna found 6 bugs to Astra's 14 with only 50% precision. The security gap was widest: 9 vs 19 verified bugs. Luna is good enough for everyday correctness bugs at that price, but not for auth or permission code on its own.
Why it matters: A controlled experiment on 50 real PRs comparing Luna and Astra on code review cost, recall, and false-positive rate — concrete and reproducible. Points off because it's a vendor self-test; the post doesn't disclose PR sources or a human reviewer baseline, so independence is u...