Codex over-engineering: 10 issues balloon to 80 in weekend test, prompting constraint strategies
2026-08-24 群聊日报
A weekend test pitted Codex, Claude Code, and Grok Bot against the same set of issues. Codex ran for 48 hours and inflated 10 issues into 80, while Claude and Grok finished in under 10 hours. Codex opened new issues even for fixes requiring only a few lines of code—70% of those new issues were meaningful, 30% imaginary. The group consensus: model capability is no longer the bottleneck; constraint engineering and taste alignment are. One member shared a checklist and Coding Convention approach to tame Sol, advocating plan review before execution. Separately, Codex reinstated a 5-hour limit for Plus users (Pro unaffected); OpenAI's internal forecast shows Go ($8/month) will capture 92% of personal subscriptions by end of 2026, with Plus dropping to 7%.
Why it matters: Real user side-by-side of three coding agents with concrete numbers — Codex over-engineered and the user canceled their $200 plan. HKR all hit. Score capped at 72 because the source is an anonymized chat digest, not a first-party benchmark, and the signal is concentrated in on...