Bebop: Accelerating RL Training via Multi-Token Prediction with Rejection Sampling
Bebop:通过带拒绝采样的多token预测加速RL训练
Alibaba researchers tackle a real bottleneck: when using multi-token prediction (MTP) for speculative decoding during RL training, acceptance rates drop sharply. They trace the cause to rising model entropy in the RL stage, which shows a clear negative linear relationship with acceptance rate. Their fix has three parts: probabilistic rejection sampling instead of greedy draft sampling, a novel end-to-end TV loss that directly optimizes multi-step acceptance rate (~10% improvement), and pre-RL MTP training that eliminates costly online updates. On Qwen3.5, 3.6, and 3.7, acceptance rates reach up to 95%, inference throughput gains hit 25%, and end-to-end async RL training speeds up to 1.8x across math reasoning, code generation, and agentic tasks.
Why it matters: Alibaba team clearly quantifies the MTP draft acceptance drop during RL training and offers concrete fixes—probabilistic rejection sampling and a TV loss. Score stays at 78 rather than 85+ because it's an arXiv preprint with no cross-source validation or deployment feedback yet.