Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation

Abstract

We conduct a controlled study of evaluation data contamination during language model pretraining at 1B and 8B scales. Contaminating both source and target text can inflate machine translation scores by up to 30 BLEU points, with larger models showing greater sensitivity; timing, frequency, and language representation also shape the effect.

Publication
Proceedings of the 42nd International Conference on Machine Learning (ICML 2025)
Muhammed Yusuf Kocyigit
Muhammed Yusuf Kocyigit
Research Scientist at Google

Research scientist working on reliable language model evaluation, data contamination, memorization, and generalization.

Related