Models Are Getting Dumber on Purpose
Small models are crushing reasoning benchmarks while their factual recall collapses. Qwen3.5 9B hallucinates 80–82% of the time on knowledge tests; Gemini 2.5 Pro hits only 53% on SimpleQA. This is a deliberate trade: labs are swapping stored facts for reasoning skill. Facts are bulky and rot; reasoning procedures compress well and don't age. The author argues a frontier-reasoning model will run on a single consumer GPU within a couple of years, but it won't know much—it will just say 'I don't know' and look things up, which may actually solve hallucination.
Why it matters: A counterintuitive industry observation backed by concrete benchmark numbers showing the reasoning-vs-recall tradeoff. Hits all three HKR axes but is commentary rather than a primary release, landing in the 78-84 band. No cross-source cluster signal, no bump.