Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
Muhammed Yusuf Kocyigit,
Eleftheria Briakou,
Daniel Deutsch,
Jiaming Luo,
Colin Cherry,
Markus Freitag
July 2025
Abstract
We conduct a controlled study of evaluation data contamination during language model pretraining at 1B and 8B scales. Contaminating both source and target text can inflate machine translation scores by up to 30 BLEU points, with larger models showing greater sensitivity; timing, frequency, and language representation also shape the effect.
Publication
Proceedings of the 42nd International Conference on Machine Learning (ICML 2025)
Research Scientist at Google
Research scientist working on reliable language model evaluation, data contamination, memorization, and generalization.