MiniMax open-sources MSA, a sparse attention method that cuts attention compute by 28.4× at 1M tokens on a 109B model
MiniMax Sparse Attention (MSA)
MiniMax published a paper introducing MSA, a blockwise sparse attention built on GQA. A lightweight index branch scores KV blocks and picks a top-k subset per GQA group, then the main branch runs exact attention only on those blocks. With a co-designed GPU kernel, a 109B-parameter multimodal model achieves 14.2× prefill and 7.6× decoding wall-clock speedups on H800 at 1M context, matching full GQA quality. Code and inference kernel are open-sourced, along with a model called MiniMax-M3. The Reddit poster is curious whether the 109B model can run on consumer GPUs; the post doesn't say if weights will be released.
Why it matters: The paper has concrete mechanisms and measured numbers, not just theory—real knowledge for inference-optimization folks. But the audience is narrow (R missed), and the low-level CUDA details raise the accessibility bar for generalist readers, so I docked 3 points, landing righ...