Skip to content
Computing Life · Share · Yage

KV Cache Hit Rate: The #1 Cost Lever for Agent Inference

Agent inference bills are dominated by prefill—re-reading the full context before every tool call—not by token generation. Spheron measured prefill at 85–95% of agent inference cost, with a 267:1 input-to-output token ratio. Raising KV cache hit rate from 0% to 90% can drop monthly GPU bills from $20K to $2K. Three engineering layers address this: compression (CompressKV retains only 3% of KV cache while keeping 97% LongBench QA performance, though FlashAttention kernels don't expose attention scores), routing (prefix-hash routing cut TTFT p90 from 92.5s to 0.54s vs. round-robin), and API-level prompt caching (Claude charges 0.1x for cached input, but Anthropic silently dropped the default TTL from 1 hour to 5 minutes in March 2026, causing 100x bill spikes). Teams running multi-turn agents should enable prompt caching before debating model choice and make cache hit rate the first dashboard metric.

Why it matters: Spheron's production measurements plus independent arXiv validation plus corroboration from Cockroach Labs and Manus make a solid case on agent inference cost structure. The 267:1 input-to-output ratio and 10x bill reduction at 90% cache hit rate are hard numbers. Scores lower...

Read the original ↗Export Markdown