Google proposes Retrieve-for-Train: shift search cost from inference to training
Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train
Google Research introduces Retrieve-for-Train (R4T), a training paradigm that moves the heavy search step of RAG from inference time to training time. During training, relevant documents for each sample are pre-fetched from an existing search index and stored in the dataset; at inference, the model uses these pre-retrieved contexts without querying the index live. The post reports 40–60% lower inference latency and 2–3× higher throughput, with quality close to real-time RAG. I'd take those numbers with a grain of salt—they come from Google's own experimental setup and may not transfer directly. The post does not disclose the base model, index size, or any open-source code.