An Open Benchmark for Testing RAG on Realistic Company-Internal Data
EnterpriseRAG-Bench released a 500k-document corpus for testing RAG on company-internal data. It simulates Redwood Inference across 9 sources and includes 500 questions over 10 retrieval failure modes. Baselines show BM25 beats vector search overall, while agentic/bash retrieval has the best completeness at higher cost and latency.
Why it matters: HKR-H/K/R all pass: the benchmark targets a real enterprise RAG pain point, with 500k docs and testable BM25-vs-vector results. Single Reddit-source benchmark release keeps it below same-day must-write.