Skip to content
最佳拍档 (BestPartners)

Breaking RLHF scaling bottlenecks: DeepMind raises data efficiency 10x with information-directed exploration

突破RLHF的规模化瓶颈 | DeepMind团队论文 | 数据利用效率极低 | 四种RLHF算法 | off-policy | 在线RLHF | 认知神经网络ENN | 信息导向探索 | 肯定性微调

A Google DeepMind team reports that online RLHF plus information-directed exploration on Gemma 9B reaches about 55% win rate with under 20k preference labels, versus about 200k for offline RLHF. The post describes four algorithms—offline, periodic, online, and information-directed exploration; online training uses batches of 64 prompts and 16 sampled responses per prompt, while the ENN head adds under 5% parameters. The key point is methodological, not that RLHF failed; the post also says results use Gemini 1.5 Pro simulated feedback, and the 1000x gain is an extrapolation toward 1M labels.

Why it matters: HKR-H/K/R all pass: the 10x label-efficiency claim is a strong hook, and the post includes concrete setup details. I kept it at 77 because this is a secondary video summary, feedback is simulated with Gemini 1.5 Pro, and the 1000x figure is an extrapolation.

Read the original ↗Export Markdown