Your model already knows the answer: how benchmark answers leak into LLMs
When benchmarks use real-world outcomes, the answer may already be in the model's training data. Elman breaks contamination into three routes: input leak (documents with the answer at test time), benchmark leak (test sets in training data), and outcome leak (public facts absorbed as general knowledge). Outcome leak is the hardest to fix—even a fresh, date-blinded test set won't stop a model from knowing a famous drug succeeded, because it learned that from papers, news, and patents during pre-training. Waiting for unresolved events is clean but impractical when drug development decisions take a decade to settle. The post surveys eight mitigation methods; most target benchmark leak, only three address outcome leak.
Why it matters: The three-leak-path breakdown is clear, and the clinical trial example is concrete. But it's a company blog from Elman with a product pitch in the background, and it frames the problem without offering a fix, so it stays at 78.