Skip to content
Hacker News front page

Can I Buy Your KV Cache?

This paper proposes letting publishers precompute a document's KV cache so AI agents can buy and load it, skipping the most compute-heavy step: prefill. On Qwen3-4B, reuse is 9–50x cheaper than prefill with zero accuracy loss—token outputs match exactly. Shipping the KV cache fails because it's nearly incompressible and egress costs more than the prefill saved. The fix: host it provider-side, like production prompt caching. Serving one 3,774-token document to 80M agents costs ~$1.5M to re-prefill but only ~$30K via reuse, a 49.7x gap. The paper frames this as an agent-native prefill CDN and leaves lossless KV compression and cross-party payments as open problems.

Why it matters: Selling precomputed KV caches is a practical idea with a 9–50× cost gap and zero accuracy loss. Held back by single-model experiments (Qwen3-4B only) and no detail on cache security or pricing in the excerpt.

Read the original ↗Export Markdown