Cerebras powers OpenAI's GPT-5.6 Sol Ultrafast at 750 tokens per second
Accelerating GPT-5.6 Sol Ultrafast
Cerebras and OpenAI previewed Ultrafast mode, running GPT-5.6 Sol on Cerebras' wafer-scale chips at up to 750 output tokens per second. On Humanity's Last Exam, it answered all 2,500 PhD-level questions in 11 hours 11 minutes—nearly 7× faster than Claude Fable 5 with comparable accuracy. On GDP-Val it delivered a 5.6× end-to-end speedup with no quality loss. The speed comes from packing 44 GB of SRAM on a single wafer, keeping model weights on-chip to avoid memory bandwidth bottlenecks. Access is limited preview for now.
Why it matters: OpenAI and Cerebras jointly unveiled Ultrafast mode for GPT-5.6 Sol, hitting 750 tok/s — a speed that pulls frontier models into real-time interaction territory. The 11-hour HLE run across 2,500 questions gives deployment teams a concrete number to work with. Not a perfect sco...