Skip to content
AI HOT (Curated Pool)

Xiaohongshu's RedKnot splits KV Cache by attention head to speed up long-context inference

小红书 RedKnot 推理引擎:将 KV Cache 按注意力头拆解实现长文本加速

RedKnot observes that 83.4%–96.8% of attention heads only care about local context, so it splits KV Cache along the head dimension and applies head-specific sparsity. Combined with sparse FFN and SegPagedAttention, it achieves up to 3.54× TTFT speedup and 7.8× higher per-card concurrency on 8×H800, cutting prefill FLOPs by 67%–79.5%. On DeepSeek-V4-Flash with 128K context, TTFT speeds up 5.16× and KV transfer drops by up to 6.3×, while accuracy stays above 95% of dense F1.

Why it matters: Xiaohongshu's RedKnot team published an inference engine optimization: splitting KV Cache per attention head and only storing full cache for heads that actually need long context. 83.4%–96.8% of heads only look locally, translating to 67%–79.5% prefill compute reduction and up...

Read the original ↗Export Markdown