Skip to content
Xinzhiyuan · WeChat

Next-ToBE Targets Short-Sighted Next-Token Prediction in LLMs at ICLR 2026

大模型自信且短视!Next-ToBE破除Next Token预测诅咒 | ICLR'26

East China Normal University and Fudan University researchers proposed Next-ToBE, a training objective that keeps standard autoregressive inference while adding a soft target over future-token windows, and the article reports the method ranked best in 35 of 36 experiments across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Llama3.1-8B-Instruct.

Why it matters: HKR-H and HKR-K pass: the mechanism and 35/36 result are specific, and next-token training is a real debate. The item stays near the featured floor because no artifact, reproduction detail, or production claim is disclosed.

Read the original ↗Export Markdown