Skip to content
r/LocalLLaMA

Prefill-as-a-Service: KV Cache of Next-Generation Models Could Go Cross-Datacenter

Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter

Moonshot says Kimi Linear makes KV cache transfer practical across datacenters, with a 20x scaled-up model showing 1.54x throughput and 64% lower P90 TTFT. The post describes prefill/decode disaggregation across datacenters and heterogeneous hardware; the cost metric and reproducibility details still require the linked arXiv paper.

Read the original ↗Export Markdown