MemStitch: Zero-copy KV cache stitching for vLLM cuts multi-agent TTFT by up to 25x
Show HN: MemStitch – Zero-copy context bridging for vLLM (25x TTFT speedup)
DaqulaLin open-sourced MemStitch, a gateway that sits in front of vLLM and stitches KV caches across requests at the memory level using PagedAttention. It skips redundant prefill in multi-agent workflows, cutting TTFT by up to 25x and saving over 40% VRAM. The post doesn't disclose the test model or GPU setup, so I'd discount the 25x claim until those details surface.
Why it matters: Multi-agent inference latency is a real pain point, and MemStitch's approach of KV cache stitching at the vLLM memory layer is more fundamental than prompt-level solutions. The 25x claim lacks disclosed model/GPU config, so the number gets a discount, but the mechanism itself ...