C2C lets LLMs talk via KV-cache, 2.5× faster and 3–5% more accurate than text
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
This ICLR'26 paper proposes Cache-to-Cache (C2C), where multiple LLMs communicate by directly exchanging KV-cache instead of generating text. A neural network projects and fuses the source model's KV-cache into the target model, with a learnable gate selecting which layers benefit. C2C beats single models by 6.4–14.2% in average accuracy, outperforms text-based communication by ~3.1–5.4%, and delivers an average 2.5× latency speedup. Code is open at thu-nics/C2C.
Why it matters: ICLR'26 paper with a clever idea: let collaborating models pass KV-cache directly instead of text. Has validation experiments and a concrete framework, so knowledge density is solid. Score capped because it's low-level optimization — not immediately actionable for most practit...