OpenAI shares first measured results for its custom inference chip, Jalapeño
The full stack behind abundant intelligence
OpenAI published the first measured results for Jalapeño, its custom inference chip. On the InferenceX benchmark running GPT‑OSS 120B, it delivered higher peak throughput per kilowatt and lower token latency than the commercial systems compared, with strong results on DeepSeek R1 and Kimi K2 as well. The post frames this as a working first-party silicon path that gives OpenAI direct control over serving economics. It also details a multi-supplier compute portfolio—Microsoft, NVIDIA, AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, SoftBank—and a self-built data center in Georgia called Project Camellia. The core argument: co-designed hardware and software lower the cost of useful intelligence, which expands usage, funds further R&D, and creates a compounding advantage.
Why it matters: OpenAI's first public benchmarks for its custom Jalapeño inference chip show better per-kW throughput and per-token latency than commercial alternatives on GPT-OSS 120B, with solid results on DeepSeek R1 and Kimi K2. This marks a key step from pure model company to full-stack ...