Skip to content
Hacker News front page

DiffusionGemma Technical Report: High-Speed Text Generation via Discrete Diffusion

DiffusionGemma Technical Report

Google's DiffusionGemma is an experimental open-weight LM that generates text by iteratively refining 256-token blocks in parallel, hitting roughly 1,500 tokens/sec on a single H100. It is fine-tuned from the MoE Gemma 4 (3.8B active, 25.2B total) using under 10% of the original training token budget. A two-stage pipeline—supervised fine-tuning for bidirectional denoising, then RL with sampler distillation—improves both quality and speed. The model sets a new Pareto frontier for speed vs. capability, retains thinking mode, multimodal inputs, and long-context support, and can still do autoregressive decoding with minor degradation.

Why it matters: DiffusionGemma applies diffusion to text generation with 256-token parallel blocks, hitting ~1,500 tok/s on a single H100 — faster than speculative decoding. Not pushing past 84 because it's still a tech report with no product path or real-world deployment data yet.

Read the original ↗Export Markdown