Pretraining progress is mostly coming from data
Dwarkesh Patel and Jerry Han trained models using year-specific open recipes and data corpora from 2019–2025. At a 1e19 FLOPs budget, data improvements delivered a 12x compute-efficiency gain versus 3.7x from model improvements, with additive effects. The authors argue model research's main value was enabling larger-scale training—MoE, FlashAttention, stability fixes—not just saving FLOPs. Small models benefit more from data quality; big models may prefer quantity over aggressive filtering, though the post doesn't confirm this at frontier scale.
Why it matters: Dwarkesh and Jerry Han ran a controlled experiment across six years of public recipes, decomposing pretraining progress into 12x from data and 3.7x from model improvements—clean conclusion backed by numbers. Not scoring higher because this is a blog post, not a peer-reviewed p...