InfiniteKV open-sourced: compresses old tokens into 104-byte searchable records on RAM or disk instead of evicting them
Open sourcing InfiniteKV: a KV cache that files old tokens as 104-byte searchable records in RAM or on disk instead of deleting them. Mistral-7B answered from token 76,747, 2.3x past its trained window. Colab demo
InfiniteKV splits the KV cache into two tiers: the latest 256 tokens stay exact in GPU memory, while older tokens are compressed into 104-byte records stored in RAM or memory-mapped disk files. For each generated token, the cache retrieves the most relevant cold records and attends over them together with the hot window—nothing is ever deleted. Mistral-7B answered a buried passkey at token 76,747, 2.3× past its trained window; at one million tokens the cold store takes roughly 3 GB versus 122 GB for float16. The author verified seven models on a 16 GB RTX 3080 laptop, reporting top-1 agreement around 0.95 and median KL divergence around 0.002 against the unmodified model. The reference implementation is pure PyTorch and slow; sliding-window and MLA models are not yet supported.
Why it matters: Open-source KV cache solution with concrete numbers and a reproducible Colab demo, directly hitting the long-context pain point for local inference. All three HKR axes hit. Score held below 85 because it's a community project (not an institutional release) and only validated o...