Skip to content
Hacker News front page

kapa.ai taught a small LLM to prune 68% of RAG context while keeping 96% recall

Pruning RAG context down to what the answer actually needs

kapa.ai inserted a cheap small LLM between retriever and generator to prune context. The pruner reads the question and all retrieved chunks together, grades them on a five-level scale, and drops 68% of context before the expensive model sees it. Net query cost fell by a third while recall held at 96%. Cutting on rerank scores failed because scores aren't calibrated across queries and can't see cross-chunk relevance. Anchor documents fixed calibration but not the underlying scores. The working design judges the set, not individual chunks.

Why it matters: kapa.ai's engineering post has real numbers and a comparison with failed approaches — not fluff. The knock is that it's one company's practice report, not a protocol update or model release, so reach is limited.

Read the original ↗Export Markdown