Skip to content
Hacker News front page

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

TryAI ran 12 models through 4 coding tasks, 5 attempts each. GPT-5.6 Sol was the most consistent—5/5 playable on the raycaster at $1.35 per run. Grok 4.5 also hit 5/5 at just $0.27, making it the value pick. Muse Spark 1.1 was erratic: 3 of 5 attempts broke, but the working ones matched Sol's quality. All raw builds and videos are linked so you can judge for yourself.

Why it matters: TryAI's 12-model coding shootout delivers pass rates, cost, and latency — the GPT-5.6 Sol vs. Grok 4.5 value gap is the headline. Capped below 84 because it's a third-party eval, not a lab release, and Muse Spark 1.1's flakiness dilutes the signal slightly.

Read the original ↗Export Markdown