Skip to content
Hacker News front page

The Economics of Open-Weight Inference

Ornn Data finds self-hosting open-weight models can cut inference cost to one-fifth of closed models. On the sparse gpt-oss-120b, an A100 undercuts an H100 at $0.12 per million output tokens. The market reflects this: five-year A100 rental contracts retain 80% of the one-month price, versus 44–60% for Hopper and Blackwell. Latency-tolerant workloads like batch eval, long-running agents, and RL can route demand to any cost-efficient hardware, extending older GPUs' earning life.

Read the original ↗Export Markdown