Skip to content
AI HOT (Curated Pool)

NVIDIA packs autoregressive, diffusion, and self-speculation decoding into one model: Nemotron-Labs-Diffusion

Nemotron-Labs-Diffusion:统一自回归、扩散与自我推测解码的三模式语言模型

NVIDIA introduces a tri-mode LM that switches between AR, diffusion, and self-speculation decoding in a single architecture. Trained with a joint AR-diffusion objective: diffusion handles lookahead planning, AR supplies left-to-right priors. In self-speculation mode, diffusion drafts and AR verifies, beating multi-token prediction (MTP) on acceptance rate and real-device efficiency. A speed-of-light analysis shows diffusion can produce up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. The family scales to 3B, 8B, and 14B parameters, covering base, instruct, and vision-language variants. The 8B model decodes 6× more tokens per forward than Qwen3-8B at comparable accuracy, yielding 4× higher throughput on SPEED-Bench with SGLang on a GB200 GPU. The post doesn't disclose open-weight plans or a release timeline.

Why it matters: NVIDIA packs AR, diffusion, and self-speculation into one model—a fresh architecture idea with concrete technical detail in the joint training objective. But it's a pure paper with no product tie-in, so resonance is weak, landing right at the featured threshold.

Read the original ↗Export Markdown