Skip to content
r/LocalLLaMA

Shard - Getting to 10× KV Cache Compression

Shard - getting to 10× KV cache compression

Shard reduces Llama-3.1-8B KV memory by about 10× at 8K context and 11× at 32K, with no measured drop on NIAH or LongBench, using PCA plus int4 quantization for K and Hadamard rotation plus vector quantization for V.

Why it matters: HKR-H/K/R all pass: the 10× KV-cache claim has a strong hook and concrete model/context/benchmark details. Reddit-only sourcing and limited validation keep it in the 78–84 band.

Read the original ↗Export Markdown