Evaluation Data Contamination in LLMs: How Do We Measure It and (When) Does It Matter?

Abstract

We introduce ConTAM, an analysis framework that grounds contamination metrics in whether models benefit from the examples marked as contaminated. A large-scale study across 13 benchmarks and seven models shows that contamination effects can be larger than release reports suggest and that metric thresholds should be model- and benchmark-specific.

Publication
arXiv preprint
Muhammed Yusuf Kocyigit
Muhammed Yusuf Kocyigit
Research Scientist at Google

Research scientist working on reliable language model evaluation, data contamination, memorization, and generalization.

Related