Skip to content
r/LocalLLaMA

Implemented TurboQuant, but results do not fully match the paper

Implemented TurboQuant and results don’t fully match paper

A Reddit user reimplemented TurboQuant and found the PROD variant reached about 95.8% correlation at 4-bit, below the paper’s 99%+ claim. They report degraded attention quality, with about 67% top-1 accuracy in a simple simulation. The key issue is correlation versus ranking preservation in KV cache quantization.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit reproduction, not a formal release. The 95.8% 4-bit correlation and ~67% top-1 result make it a low featured item.

Read the original ↗Export Markdown