FastDMS: 6.4X KV-cache compression running faster than vLLM BF16/FP8
FastDMS released an MIT implementation that cuts KV memory to 1/5–1/8 of vLLM BF16 at 8K context. A Llama-3.2-1B replication reports PPL 9.200 with 6.4x compression; Qwen3-8B c=1 drops KV from 1.406 GiB to 0.184 GiB. The key detail is physical reclamation of evicted slots, not just nominal KV-byte reduction.
Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, with compression, PPL, KV GiB deltas, and physical slot reclamation. Reddit/open-source sourcing keeps it in 78–84, below P1.