Skip to content
Google Research Blog

Parametric factuality errors are mostly recall failures, not knowledge gaps

Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

Google Research splits factual errors into two types: knowledge not in the model (empty shelves) and knowledge the model learned but fails to retrieve (lost keys). Across 4 models and 6 datasets, at least 70% of errors are retrieval failures—the correct answer appeared in training but wasn't surfaced at inference. For Gemini 2.5 Pro, over 90% of factual mistakes fall into this bucket. The team used a probing method called SIR, feeding training data to check whether the model's internal state can activate the right answer. The takeaway: improving retrieval beats stuffing in more knowledge.

Why it matters: Google Research uses SIR probing to split factual errors into 'never learned' vs 'can't recall,' finding ≥70% are recall failures across 4 models and 6 datasets. Directly useful for practitioners, but missing breakdowns by model scale keep it from 85+.

Read the original ↗Export Markdown