Three separate stories, one shared distortion: prompt caching makes headline numbers look bigger than they are.
OpenRouter says agents burn 7.3T tokens weekly, 5.2× human usage. Sounds dramatic, but 70–85% are cached reads billed at roughly 10% of full price. The real bill is closer to 2×. And this is one platform—OpenRouter claims only ~1% of global inference. Their human-vs-agent classification uses seven weighted signals with undisclosed weights. One app, Hermes Agent, accounts for 20% of agent traffic alone.
GitHub's Knowledge Compressor halves doc length and claims breakeven at 2,000 reuses. That's at full price. With 50% cache hit rate, breakeven jumps to 3,640; at 90%, it's 10,500; all-cached hits 20,000. Median expectation is above 5,000. The prototype isn't open-source, and Q&A-based validation misses connective knowledge loss.
OpenAI's Jalapeño chip beats GB200/GB300 on 8k/1k fixed-length benchmarks—1.5–1.9× throughput per watt. But SemiAnalysis notes that real agent workloads are dominated by repeated long-context reads, exactly the part covered by caching discounts. AgentX scores for realistic workloads are still missing. Jalapeño is an engineering sample shipping at tiny scale in late 2026, while Rubin is already shipping.
Same pattern across all three: headline numbers look impressive, but once you factor in caching discounts, the conclusions shift. Ask whether the cache discount has been applied before you quote the number.