Parametric factuality errors are mostly recall failures, not knowledge gaps
Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
Google Research splits factual errors into two types: knowledge not in the model (empty shelves) and knowledge the model learned but fails to retrieve (lost keys). Across 4 models and 6 datasets, at least 70% of errors are retrieval failures—the correct answer appeared in training but wasn't surfaced at inference. For Gemini 2.5 Pro, over 90% of factual mistakes fall into this bucket. The team used a probing method called SIR, feeding training data to check whether the model's internal state can activate the right answer. The takeaway: improving retrieval beats stuffing in more knowledge.
Why it matters: Google Research uses SIR probing to split factual errors into 'never learned' vs 'can't recall,' finding ≥70% are recall failures across 4 models and 6 datasets. Directly useful for practitioners, but missing breakdowns by model scale keep it from 85+.