Skip to content
Computing Life · Share · Yage

The Four-Year History of Reasoning Models: The Quiet Thread Before the Breakthrough

Reasoning models didn't appear overnight in 2024. Chain-of-thought prompting, STaR self-training, process reward models, and test-time compute scaling laws all predate o1. What o1 actually changed was productization: turning reasoning into a billable, schedulable resource and opening a second axis for scaling. DeepSeek R1 made the know-how public, triggering industry-wide convergence within five months. But the most hyped part—pure RL spontaneously creating reasoning—is the weakest claim. Independent studies show base models already contain reasoning fragments; RL merely amplifies their frequency. The real lesson: distinguish the birth of a capability from its packaging.

Why it matters: A well-researched long-read that traces reasoning model lineage with specific papers and timelines, arguing the real o1 watershed was productizing reasoning as billable compute, not inventing it. HKR all hit, but it's a synthesis piece rather than a scoop — lands at 78, the fe...

Read the original ↗Export Markdown