Alibaba RTPurboV2 uses hundreds of training steps for 10x sparse attention
阿里RTPurboV2:原生Transformer再次崛起,百步训练实现10倍稀疏注意
Alibaba’s RTP team released RTPurboV2, replacing 85% of attention heads with SWA and compressing the remaining 15% retrieval heads using low-rank projection, clustering, and dynamic top-p; the adaptation uses about 600 training steps and roughly 1M label tokens, with reported Prefill speedup up to 9.36x.
Why it matters: HKR-H/K/R all pass: the hook is 100-step training and 10x sparse attention, with concrete mechanisms and 9.36x Prefill speedup. Strong engineering signal from Alibaba, but not a flagship model release or major product launch.