Jan Iłowski ran the numbers using Artificial Analysis benchmark data, and the takeaway is blunt: picking models by per-token sticker price will likely cost you more for worse results.
Two things break the comparison. First, tokenizers differ across labs—Anthropic recently tweaked theirs, adding 30% more tokens for the same text, which is effectively a hidden price hike. Second, reasoning tokens—the model's internal "thinking"—are billed at the same rate as output tokens but vary wildly between models. That's where most of your real spend goes.
The table has two standout rows. GPT-5.5 xhigh lists at $5/$30 per million tokens, Claude Opus 4.8 max at $5/$25—looks close. But per completed benchmark task, GPT-5.5 costs $0.99 while Opus 4.8 runs $1.78. DeepSeek V4 Pro max is the extreme outlier: $0.435/$0.87 sticker, roughly $0.04–$0.05 per task, though its benchmark score of 44 trails GPT-5.5's 55 by a clear margin.
The author's take on Claude Sonnet 5 max is worth noting: $3/$15 sticker looks cheaper than Opus, but $2.29 per task is actually higher—token efficiency is noticeably worse. He half-jokingly calls it an Anthropic psy-op to lure people in with low sticker prices. I'd say that's a bit conspiratorial, but the data does show Sonnet 5 underperforming on cost-efficiency in this benchmark.
This piece matters because it pulls the pricing conversation back to actual consumption. If you're picking an API provider, run a few of your own typical tasks and measure real cost—that'll tell you more than any pricing page.