Jon Clegg's test is refreshingly simple: one prompt, one HTML file, scored automatically by Opus 5.5. Claude Opus 5.5 hit 99/100 with a 10.8 KB page that nailed the arcade audio—full two-phrase intro with a bass line, continuous siren, proper death sound—for $1.99 in 9 minutes. Fable 5.1 scored 96 but cost $5.87, which is a worse deal. GPT-5.6-sol got 90 for just $0.72 in under 5 minutes; if you don't care about audio fidelity, that's the budget pick.
Don't read this as a raw model coding benchmark. It measures model-plus-toolchain combos—Claude Code, Grok Build, and other harnesses that can iterate and fix bugs, not a single-shot generation from a bare model. Sonnet 5.5 ran in Phase 2, which suggests the harness did something extra, but the post doesn't explain what Phase 2 means. Also, Opus 5.5 is both contestant and judge; the scoring rubric is public, but I'd discount small score differences when the same model is doing the grading.
The useful bit is the cost comparison: GPT-5.6-sol at $0.72 is less than half of Opus's $1.99 for a playable prototype, though the audio and polish take a hit. Pick based on whether speed or finish matters more.