This piece lands because of who's talking. Zach Mueller runs developer ecosystem at Lambda, a GPU cloud company. He built a home GPU rack to run open-weight models, then publicly admitted it doesn't save money—his electricity bill might double, and he's stopped trying to justify it financially. The return is all skill investment. When someone who sells GPU cloud says self-hosting isn't cheaper, it carries more weight than any benchmark chart.
The article breaks the decision into three ledgers, and the logic is clean. First ledger is cost: two H100s need roughly 2 billion tokens per month to break even with mid-tier cloud APIs. That throughput is way beyond what most small teams generate. The real trap is that utilization and latency fight each other—cram more requests into the same GPUs to lower per-token cost, and time-to-first-token jumps from 45ms to 740ms. You can't have both cheap and snappy. Second ledger is data compliance: 31% of enterprises rank data privacy as their top concern, but OpenAI and Anthropic now offer zero-retention enterprise agreements. A lot of compliance needs can be solved with a contract, no GPU rack required. Third ledger is capability: a single LoRA fine-tune starts at $300, and vertically tuned small models can beat general-purpose large models on specific tasks—Hamel Husain's Honeycomb and ReChat case studies back this up.
The article maps three tiers: rent tokens, rent hardware, own hardware. Zach's practical advice: rent spot instances for two to three weeks first, track actual token usage, then decide. A single RTX 4090 costs about $104/month all-in, and you'd need 8 million tokens monthly to match GPT-4o-level API pricing. I'm bookmarking this for the next time someone asks me whether they should self-host.