Google releases DiffusionGemma, a diffusion-based text model that generates 4× faster than autoregressive models
Mythos阴影里谷歌悄悄发模型,速度暴涨4倍
Google open-sourced DiffusionGemma, a 26B MoE diffusion text model that activates only 3.8B parameters at inference and fits in 18GB VRAM after quantization. It denoises 256 tokens in parallel—like a printing press instead of a typewriter—hitting 1,000+ tokens/s on an H100 and 700+ on an RTX 5090, roughly 4× faster than a comparable autoregressive model. Bidirectional attention enables real-time self-correction; after fine-tuning, Sudoku accuracy jumped from 0% to 80%. Quality still trails Gemma 4, and Google positions it as an experimental “racehorse” for speed-sensitive local use. Released under Apache 2.0, weights available on Hugging Face.
Why it matters: Google open-sourced DiffusionGemma, applying diffusion models to text generation with 256 tokens denoised simultaneously, roughly 4x faster than comparable autoregressive models. Score isn't higher because only speed numbers are out—generation quality and downstream task perfo...