The sample efficiency black hole: AI models need far more data than humans to learn
The sample efficiency black hole
Dwarkesh Patel argues that recent AI progress comes from more and better data, not better sample efficiency. RL is framed as synthetic data generation: spend compute to find good rollouts, then train the model to predict them. Each skill requires hundreds of human experts writing examples and rubrics, fueling a data-labeling industry earning billions annually. A human sees ~200M tokens by adulthood; frontier models train on tens to hundreds of trillions—a nearly million-fold gap. A person learns to teleoperate a robot in hours, while self-driving models need 3–4 orders of magnitude more data than a teen learning to drive. Open models lag closed ones by only 4 months because data is easy to distill from public APIs, unlike architecture tricks. The post does not propose a fix for sample efficiency.
Why it matters: Dwarkesh reframes RL as synthetic data generation and contrasts human vs. model data consumption with concrete numbers—high information density. Score capped at 78 because it's an opinion piece, not a primary release, and some arguments lean on analogy rather than experimental...