Stop Thinking of LLMs as Next-Token Predictors
Calling LLMs 'next-token predictors' is technically true but misses the point: post-training (especially RLVR) lets models explore new sequences and learn from rewards, not just imitate existing text. The author uses a chess analogy: one system predicts grandmaster moves from a database, another explores all possible games and picks the winning move—the latter is not a 'next-move predictor.' The post doesn't name specific models but explains how RLHF and RLVR shift models from imitation to simulation and discovery.