Evaluation Data Contamination in LLMs: How Do We Measure It and (When) Does It Matter?
Aaditya K. Singh,
Muhammed Yusuf Kocyigit,
Andrew Poulton,
David Esiobu,
Maria Lomeli,
Gergely Szilvasy,
Dieuwke Hupkes
November 2024
Abstract
We introduce ConTAM, an analysis framework that grounds contamination metrics in whether models benefit from the examples marked as contaminated. A large-scale study across 13 benchmarks and seven models shows that contamination effects can be larger than release reports suggest and that metric thresholds should be model- and benchmark-specific.
Publication
arXiv preprint
Research Scientist at Google
Research scientist working on reliable language model evaluation, data contamination, memorization, and generalization.