Skip to content
Trending storyPast story

The Economics of Open-Weight Inference

1 report1 sourceupdated 7 days ago

What happened

Summary

Ornn Data 这份报告算了一笔账:自己租 GPU 跑开源模型,生成每百万 token 的成本可以压到 0.12 到 0.35 美元,差不多是调用闭源模型价格的五分之一。有意思的是,在跑 gpt-oss-120b 这种稀疏模型(实际干活只用 51 亿参数)时,老款 A100 反而比 H100 更省钱,满负荷下每百万 token 只要 0.12 美元...

Coverage

Follow the reports to see the story from different sides.

Sep 22
  1. Hacker News front page
    The Economics of Open-Weight Inference

    Ornn Data finds self-hosting open-weight models can cut inference cost to one-fifth of closed models. On the sparse gpt-oss-120b, an A100 undercuts an H100 at $0.12 per million output tokens. The market reflects this: five-year A100 rental contracts retain 80% of the one-month price, versus 44–60% for Hopper and Blackwell. Latency-tolerant workloads like batch eval, long-running agents, and RL can route demand to any cost-efficient hardware, extending older GPUs' earning life.