Training a single transformer layer can match full-parameter RL training
Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train
This paper shows that most gains from RL post-training concentrate in a few middle layers. Training a single transformer layer can recover or even exceed full-parameter RL results. The pattern holds across seven model sizes from Qwen3 and Qwen2.5, three RL algorithms (GRPO, GiGPO, Dr. GRPO), and tasks including math reasoning, code generation, and agentic decision-making. Middle layers dominate; input and output layers contribute little. Layer contribution rankings are strongly correlated across datasets, tasks, model families, and algorithms.
Why it matters: The paper finds most RL post-training gains concentrate in middle layers—training just one layer recovers full-parameter performance. That's directly useful for teams doing RL fine-tuning. Score isn't higher because it's an arXiv preprint without peer review yet; I'm discounti...