Evaluated a RAG Chatbot: The Most Expensive Model Was the Worst Performer
Evaluated a RAG chatbot and the most expensive model was the worst performer. Notes on what actually moved the needle.
The author evaluated a customer-support RAG bot and raised the quality score from 6.62 to 7.88 while cutting per-session cost from $0.002420 to $0.000509, using retrieval logging, LLM-as-judge scoring, chunk deduplication, stricter grounding, and a five-model sweep.
Why it matters: HKR-H/K/R all pass: counterintuitive model ranking, concrete quality and cost deltas, and direct RAG production relevance. Reddit source authority keeps it near the featured floor despite the first-person experiment signal.