Price per 1M tokens is a misleading way to compare models
Price per 1M tokens is meaningless
Jan Iłowski argues that per-token pricing hides real costs. Using Artificial Analysis benchmark data, he shows GPT-5.5 xhigh costs nearly half as much per completed task as Claude Opus 4.8 max ($0.99 vs $1.78) despite higher sticker prices. Two factors break the comparison: tokenizers differ across labs—Anthropic's recent change added 30% more tokens for the same text—and hidden reasoning tokens dominate real-world spend. DeepSeek V4 Pro max is the extreme outlier at ~$0.04–$0.05 per task. Claude Fable 5 tops the benchmark but costs $3.25 per task, over 3× GPT-5.5. The takeaway: ignore cost per task and you'll likely pay more for worse results.
Why it matters: Has concrete benchmark data and cost comparison, not just opinion; the 30% hidden price hike from tokenizer changes is practically useful for practitioners. Deduction because it's a personal blog, not an official release, and only the opening is provided—full argument strength...