This one's worth opening because Google didn't train a diffusion model from scratch—they converted the fully post-trained Gemma 4 26B-A4B weights directly, and open-sourced the result. The approach is efficient: adaptation training plus sampler distillation used under 10% of the original training token budget, letting the model denoise a 256-token canvas in parallel. On a single H100 at FP8 with batch size 1, pure decode hits 1,456 tok/s, over 7× the original AR model.
The speedup comes from a real hardware bottleneck. At low concurrency, AR models waste most of their time reading weights and KV cache from memory for each single-token forward pass. Diffusion processes 256 positions per forward pass, cutting the number of serial rounds and the memory-bandwidth overhead.
The trade-off shows up in math. AIME 2026 drops from 88.3 to 69.1, and MRCR 128K from 44.1 to 32.0. The key ablation: switching to AR fallback mode recovers AIME to 84.2, so the base knowledge survived—the diffusion generation step itself caused the quality loss.
Two deployment catches. TTFT jumps from 53 ms to 489 ms, which users will feel. And at high concurrency, AR total throughput overtakes diffusion, so this is best for low-concurrency, long-output workloads where per-token latency matters less.
Don't read this as "diffusion replaces autoregressive." The more accurate take: Google shipped a usable diffusion inference option with open weights and vLLM support that engineering teams can test today. But the paper doesn't disclose absolute token count, FLOPs, or GPU hours for the adaptation phase, so the full cost picture is still incomplete.