Skip to content
Hacker News front page

Why LLMs fail at tabular prediction: experiments rule out four hypotheses, point to dimensionality

Why Large Language Models Fail at Tabular Prediction

Marta Garnelo and Wojciech Czarnecki test a frontier LLM on tabular prediction in its purest inference mode—no tools, no fine-tuning, one generation pass. They falsify four common explanations: noisy data, CSV formatting, numeric tokenization, and number of test points per query. Dimensionality is the decisive factor. Across 31 datasets, the LLM's accuracy drops as dimensionality grows, while nine classical baselines stay flat or improve. In 2D, the LLM behaves like a local distance-based method (up to 91.6% grid agreement); in higher dimensions, no classical model—even with tuned noise—reproduces its predictions. The internal mechanism remains open, but the result explains why LLMs keep losing to decades-old baselines on tables.

Why it matters: The paper systematically tests five common explanations in the purest inference regime and identifies dimensionality as the key bottleneck—a concrete, experiment-backed finding. But it's a pure academic paper with no product or tool release, and tabular prediction is a vertica...

Read the original ↗Export Markdown