kapa.ai taught a small LLM to prune 68% of RAG context while keeping 96% recall
Pruning RAG context down to what the answer actually needs
kapa.ai inserted a cheap small LLM between retriever and generator to prune context. The pruner reads the question and all retrieved chunks together, grades them on a five-level scale, and drops 68% of context before the expensive model sees it. Net query cost fell by a third while recall held at 96%. Cutting on rerank scores failed because scores aren't calibrated across queries and can't see cross-chunk relevance. Anchor documents fixed calibration but not the underlying scores. The working design judges the set, not individual chunks.
Why it matters: kapa.ai's engineering post has real numbers and a comparison with failed approaches — not fluff. The knock is that it's one company's practice report, not a protocol update or model release, so reach is limited.