Pac-Bench: One-shot Pac-Man benchmark, Claude Opus 5.5 scores 99/100
Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?
Jon Clegg built a Pac-Man benchmark: one prompt, one HTML page, scored automatically by Opus 5.5. Claude Opus 5.5 hit 99/100 via Claude Code at $1.99, generating a 10.8 KB page in 9 minutes with near-arcade audio. Claude Fable 5.1 scored 96 but cost $5.87. Grok 4.7 and GPT-5.6-sol scored 94 and 90; the latter cost just $0.72 in under 5 minutes. Scoring covers controls, ghost behavior, stuck detection, maze layout, and sound. The post doesn't explain why some models ran Phase 2 or how much the harness affects scores. Worth noting: this measures model-plus-toolchain combos, not bare model capability.
Why it matters: A 30-model Pac-Man benchmark with Claude Opus 5.5 hitting 99/100 via Claude Code at $1.99 is solid signal. Capped at 78 because it's an individual project, not an official release, so authority is limited despite strong HKR.