Skip to content
Hacker News front page

Tokens Too Cheap to Meter

jyn argues with multiple charts that AI inference cost is dropping by orders of magnitude each year. GPU efficiency doubles roughly every two years, and per-task model cost in 2026 is two orders of magnitude cheaper than end of 2025. Inference engines like vLLM add 10%–50% throughput gains annually. The author expects LLMs to become computing infrastructure within 1–2 years, and frontier-quality local models on commodity hardware in 3–6 years. The post doesn't cite specific dollar figures, but the trend lines are stark.

Why it matters: A data-backed cost trend analysis, not vague 'AI is getting cheaper' talk. Charts three decline curves — GPU efficiency, inference engine optimization, local deployment — and engages with Jevons paradox and ROI questions. Docked slightly for being a personal blog rather than i...

Read the original ↗Export Markdown