Do Transformers Need Three Projections? Systematic Study of QKV Variants
Do transformers need three projections? Systematic study of QKV variants
Ali Kayyam and coauthors evaluate three QKV projection-sharing variants across synthetic, vision, and language-modeling settings, including 300M and 1.2B parameter models trained on 10B tokens; Q-K=V halves the KV cache with a 3.1% perplexity degradation, while Q-K=V plus MQA reduces cache use by 96.9%.
Why it matters: HKR-H/K/R all pass: the title challenges a core architecture default, the paper gives testable 300M/1.2B and 10B-token results, and KV-cache cuts map to inference cost. It remains an arXiv architecture study, so 78–84 fits.