This one's worth opening because OpenAI finally showed real silicon benchmarks, not slides. Jalapeño ran GPT-OSS 120B, DeepSeek R1, and Kimi K2 on the public InferenceX benchmark, delivering higher peak throughput per kilowatt and lower per-token latency than the commercial systems they compared against. That gives them a working first-party path to control serving economics directly.
I'd discount it a bit: this is an OpenAI blog post, and they didn't name the commercial systems they beat. H200? B200? Unknown configs. But the signal is real—Jalapeño is running production-scale models across model families, not just tuned for their own.
The bigger picture is the multi-supplier portfolio they laid out alongside it: Microsoft, NVIDIA, AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, SoftBank, plus their self-built Georgia data center Project Camellia. The logic is straightforward: premium systems for training, custom silicon to drive down inference cost, different providers for different workloads, keeping pricing discipline and roadmap flexibility.
Sarah Friar frames this as a compounding flywheel—co-designed hardware and software lower the cost of intelligence, which expands usage, which funds more R&D. Whether that flywheel actually spins depends on real deployment scale and the cost gap Jalapeño can sustain. Right now we've got benchmark numbers, not latency distributions or unit economics at scale. Hold the excitement.