OpenAI adds 'Ultrafast' mode to GPT-5.6 Sol, pushing inference to 14x speed
OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speed
OpenAI launched a preview of Ultrafast, a mode for its top model GPT-5.6 Sol that hits 750 tokens/sec — 14x the standard speed. The company says this avoids the old trade-off of switching to a smaller model for real-time use. It runs on Cerebras chips and is limited to a small customer group for now, with broader access planned. Anthropic's Claude has a fast mode but doesn't match this throughput. Target workflows include incident response, customer support, financial analysis, and e-commerce. The post doesn't disclose pricing, latency details, or a general release date.
Why it matters: 14x speed on GPT-5.6 Sol is a real inference win with direct agent implications. Held back from higher score because it's a limited preview with no GA timeline and ties to Cerebras hardware — general availability is unproven.