Grok 4.5, GPT-5.5, and Claude build the same apps: speed, cost, and quality compared
We made Grok 4.5, GPT-5.5, and Claude build the same apps
TryAI gave Grok 4.5, GPT-5.5, Claude Opus 4.8, and Fable 5 the same three app prompts and measured latency and cost. Claude models nailed the 3D Rubik's cube first try; Grok 4.5 needed its one allowed retry after a blank render, and GPT-5.5 only drew a single dark face. All four shipped a working particle sandbox and a playable Breakout game. Grok 4.5 led on speed: 0.44s first token, ~110 tok/s throughput, and the cheapest per reply. Fable 5 was slowest and priciest. The post doesn't disclose parameter counts or training details.
Why it matters: First-hand coding shootout with concrete failure cases and cost data, not just benchmark scores. Score isn't higher because TryAI isn't a tier-1 evaluator and the excerpt only gives a summary — full data requires clicking through.