Skip to content
Trending storyPast story

Google proposes Retrieve-for-Train: shift search cost from inference to training

1 report1 sourceupdated 9 days ago

What happened

Summary

Google Research 发了一篇博客,介绍一种叫 Retrieve-for-Train(R4T)的做法。简单说,就是把 RAG(外挂资料库)里最耗时的实时搜索步骤,从推理时搬到训练时。训练时,系统会提前从已有的搜索索引里,给每个训练样本找好相关文档,直接存进数据集;推理时模型直接用这些预存好的上下文,不用再现场查索引。博客给出的数据是推理延迟降...

Coverage

Follow the reports to see the story from different sides.

Sep 16
  1. Google Research Blog
    Google proposes Retrieve-for-Train: shift search cost from inference to training

    Google Research introduces Retrieve-for-Train (R4T), a training paradigm that moves the heavy search step of RAG from inference time to training time. During training, relevant documents for each sample are pre-fetched from an existing search index and stored in the dataset; at inference, the model uses these pre-retrieved contexts without querying the index live. The post reports 40–60% lower inference latency and 2–3× higher throughput, with quality close to real-time RAG. I'd take those numbers with a grain of salt—they come from Google's own experimental setup and may not transfer directly. The post does not disclose the base model, index size, or any open-source code.