This piece is worth opening because it dismantles the viral 'OpenAI collapse' article that's been circulating. That article's most explosive claim—'French is 50-100x more compute-efficient than English'—came from a blog comment posted the same day, not a paper. The one peer-reviewed study on this topic found the opposite: French is more compute-hungry due to its complex morphology.
The scaling law story itself is real and more useful than any fraud narrative. OpenAI's 2020 Kaplan paper concluded 'prioritize parameters' due to experimental biases, shaping GPT-3. DeepMind's 2022 Chinchilla corrected the ratio to 20 tokens per parameter, showing smaller models with more data outperform. Two 2024 replication studies confirmed that fixing Kaplan's setup reproduces Chinchilla's result—no fraud, just calibration.
Since 2023, the industry has been deliberately deviating from Chinchilla. Meta fed 15 trillion tokens to the 8B Llama 3, nearly 1,900 tokens per parameter—over 90x Chinchilla's ratio. The reason: the optimization target shifted from training cost alone to total cost of training plus inference. Smaller models are cheaper to run, so over-training them pays off at scale.
Tsinghua's Densing Law quantifies this: the parameter count needed for equal capability halves roughly every 3.5 months. But there's a floor—each parameter stores only ~2 bits of knowledge, so tiny models can't hold many facts. The likely future is a split: everyday tasks go to shrinking models, frontier capabilities stay with the big ones.