Sumi: Open Uniform Diffusion Language Model from Scratch
Sumi:从头训练的7B开源均匀扩散语言模型
Sumi is the first uniform diffusion language model pretrained from scratch at 7B parameters on 1.5T tokens, with weights and full training recipe released openly. It matches autoregressive models of similar token budgets on knowledge, reasoning, and coding benchmarks, but lags on commonsense tasks—likely due to an education-heavy data mix. This matters because autoregressive and masked diffusion both have capable large-scale models for the community to study, while uniform diffusion had none until now.
Why it matters: First 7B diffusion LM trained from scratch with 1.5T tokens, fully open weights and recipe, competitive with autoregressive models at equal compute. H and K are solid, but diffusion LMs aren't in mainstream workflows yet, so R is weak — lands at 78, the featured threshold.