OpenAI's first inference chip Jalapeño shows lower latency and higher throughput per watt
Jalapeño’s first results show industry-leading speed and efficiency in AI inference
OpenAI shared first measured results for Jalapeño, its custom inference chip. Across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, Jalapeño delivered 1.5–1.9× more throughput per watt at peak and 1.7–3.6× lower end-to-end latency than the comparison systems. For interactive workloads the lead widened to 2.1–4.1×. OpenAI says the chip achieves both higher throughput and lower latency without the usual tradeoff. The chip design was accelerated by OpenAI's own models. The post does not name the comparison hardware, process node, production timeline, or pricing.
Why it matters: OpenAI's first public silicon benchmark, with head-to-head numbers against three major open-weight models. The per-watt throughput and interactive latency multiples are concrete. This is the paper-to-silicon inflection point for their hardware roadmap, with real implications f...