Skip to content
AI HOT (Curated Pool)

OpenAI and Broadcom unveil Jalapeño, their first LLM-optimized inference chip

OpenAI 与 Broadcom 发布面向 LLM 推理的定制芯片 Jalapeño

OpenAI and Broadcom announced Jalapeño, a chip built from scratch for LLM inference. It went from design to production in nine months, with OpenAI's own models helping accelerate the tape-out. Early testing shows substantially better performance per watt than current state-of-the-art, though detailed benchmarks won't arrive for a few months. The chip is already running GPT‑5.3‑Codex‑Spark at production frequency and power in the lab. OpenAI says Jalapeño is not a repurposed general accelerator—it was architected around the serving patterns of ChatGPT, Codex, and future agentic products, aiming to match top training chips on throughput while approaching specialized inference systems on latency. Gigawatt-scale deployment with Microsoft and other partners begins in 2026.

Why it matters: OpenAI's first custom inference chip, taped out in 9 months and already running GPT-5.3-Codex-Spark with claimed perf/watt gains. This is a major vertical integration move, directly comparable to Google's TPU path. Score held below 90 because concrete benchmarks are months awa...

Read the original ↗Export Markdown