Skip to content
Hacker News front page

C2C lets LLMs talk via KV-cache, 2.5× faster and 3–5% more accurate than text

Cache-to-Cache: Direct Semantic Communication Between Large Language Models

This ICLR'26 paper proposes Cache-to-Cache (C2C), where multiple LLMs communicate by directly exchanging KV-cache instead of generating text. A neural network projects and fuses the source model's KV-cache into the target model, with a learnable gate selecting which layers benefit. C2C beats single models by 6.4–14.2% in average accuracy, outperforms text-based communication by ~3.1–5.4%, and delivers an average 2.5× latency speedup. Code is open at thu-nics/C2C.

Why it matters: ICLR'26 paper with a clever idea: let collaborating models pass KV-cache directly instead of text. Has validation experiments and a concrete framework, so knowledge density is solid. Score capped because it's low-level optimization — not immediately actionable for most practit...

Read the original ↗Export Markdown