We conduct a controlled study of evaluation data contamination during language model pretraining at 1B and 8B scales. Contaminating both source and target text can inflate machine translation scores by up to 30 BLEU points, with larger models showing greater sensitivity; timing, frequency, and language representation also shape the effect.
@inproceedings{kocyigit2025overestimation,title={Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation},author={Kocyigit, Muhammed Yusuf and Briakou, Eleftheria and Deutsch, Daniel and Luo, Jiaming and Cherry, Colin and Freitag, Markus},booktitle={Proceedings of the 42nd International Conference on Machine Learning},volume={267},pages={31105--31132},year={2025},}
2024
Evaluation Data Contamination in LLMs: How Do We Measure It and (When) Does It Matter?
Aaditya K. Singh, Muhammed Yusuf Kocyigit, Andrew Poulton, David Esiobu, Maria Lomeli, Gergely Szilvasy, and Dieuwke Hupkes
We introduce ConTAM, an analysis framework that grounds contamination metrics in whether models benefit from the examples marked as contaminated. A large-scale study across 13 benchmarks and seven models shows that contamination effects can be larger than release reports suggest and that metric thresholds should be model- and benchmark-specific.
@article{singh2024contamination,title={Evaluation Data Contamination in LLMs: How Do We Measure It and (When) Does It Matter?},author={Singh, Aaditya K. and Kocyigit, Muhammed Yusuf and Poulton, Andrew and Esiobu, David and Lomeli, Maria and Szilvasy, Gergely and Hupkes, Dieuwke},journal={arXiv preprint arXiv:2411.03923},year={2024},}
2023
A Novel Method for Analysing Racial Bias: Collection of Person Level References
Muhammed Yusuf Kocyigit, Anietie Andy, and Derry Wijaya
We introduce person-based filtering for studying differences in how demographic groups are represented over time. Applying the method to African American and White American figures in Google Books from 1850 to 2000 reveals historical changes in representation while reducing the selection bias of fixed group-reference phrases.
@article{kocyigit2023racial,title={A Novel Method for Analysing Racial Bias: Collection of Person Level References},author={Kocyigit, Muhammed Yusuf and Andy, Anietie and Wijaya, Derry},journal={arXiv preprint arXiv:2310.15847},year={2023},}
Western, Religious or Spiritual: An Evaluation of Moral Justification in Large Language Models
We evaluate the moral perspectives used by large language models when suggesting or judging actions. The models favor Western-tradition justifications and may over-align with religious framing, sometimes failing to identify an immoral action when it is presented as religious.
@inproceedings{kucuk2023western,title={Western, Religious or Spiritual: An Evaluation of Moral Justification in Large Language Models},author={Kucuk, Eyup Engin and Kocyigit, Muhammed Yusuf},booktitle={NeurIPS 2023 MP2 Workshop},year={2023},}
2022
AugCSE: Contrastive Sentence Embedding with Diverse Augmentations
Zilu Tang, Muhammed Yusuf Kocyigit, and Derry Tanti Wijaya
In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing, 2022
We present AugCSE, a unified framework for using diverse data augmentations to produce stronger general-purpose sentence embeddings. An adversarial objective reconciles augmentations that would otherwise create conflicting contrastive signals, improving downstream transfer while remaining competitive on semantic textual similarity.
@inproceedings{tang2022augcse,title={AugCSE: Contrastive Sentence Embedding with Diverse Augmentations},author={Tang, Zilu and Kocyigit, Muhammed Yusuf and Wijaya, Derry Tanti},booktitle={Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing},pages={375--398},year={2022},doi={10.18653/v1/2022.aacl-main.30},}
Challenges in Measuring Bias via Open-Ended Language Generation
Afra Feyza Akyurek, Muhammed Yusuf Kocyigit, Sejin Paik, and Derry Tanti Wijaya
In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing, 2022
We analyze how prompt sets, metrics, automatic tools, and sampling strategies affect measurements of bias in open-ended language generation. The study shows that different experimental choices can produce contradictory results and offers reporting recommendations for a more complete view of model behavior.
@inproceedings{akyurek2022challenges,title={Challenges in Measuring Bias via Open-Ended Language Generation},author={Akyurek, Afra Feyza and Kocyigit, Muhammed Yusuf and Paik, Sejin and Wijaya, Derry Tanti},booktitle={Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing},pages={76--83},year={2022},doi={10.18653/v1/2022.gebnlp-1.9},}
On Measuring Social Biases in Prompt-Based Multi-Task Learning
Afra Feyza Akyurek, Sejin Paik, Muhammed Yusuf Kocyigit, Seda Akbiyik, Serife Leman Runyun, and Derry Tanti Wijaya
In Findings of the Association for Computational Linguistics: NAACL 2022, 2022
We study whether semantically equivalent input formats change the social biases expressed by the prompt-trained T0 model. Across question-answer and premise-hypothesis formulations, the model exhibits substantially different bias, showing that evaluation results depend on how an input is encoded.
@inproceedings{akyurek2022measuring,title={On Measuring Social Biases in Prompt-Based Multi-Task Learning},author={Akyurek, Afra Feyza and Paik, Sejin and Kocyigit, Muhammed Yusuf and Akbiyik, Seda and Runyun, Serife Leman and Wijaya, Derry Tanti},booktitle={Findings of the Association for Computational Linguistics: NAACL 2022},pages={551--564},year={2022},doi={10.18653/v1/2022.findings-naacl.42},}
Better Quality Estimation for Low Resource Corpus Mining
Muhammed Kocyigit, Jiho Lee, and Derry Wijaya
In Findings of the Association for Computational Linguistics: ACL 2022, 2022
We improve machine translation quality estimation through multitask learning, data augmentation, and contrastive learning. The resulting models are substantially more robust in parallel corpus mining and offer a viable low-resource approach trained with roughly one thousand times less parallel data than high-resource baselines.
@inproceedings{kocyigit2022better,title={Better Quality Estimation for Low Resource Corpus Mining},author={Kocyigit, Muhammed and Lee, Jiho and Wijaya, Derry},booktitle={Findings of the Association for Computational Linguistics: ACL 2022},pages={533--543},year={2022},doi={10.18653/v1/2022.findings-acl.45},}
2020
NUBIA: NeUral Based Interchangeability Assessor for Text Generation
Hassan Kane, Muhammed Yusuf Kocyigit, Ali Abdalla, Pelkins Ajanoh, and Mohamed Coulibali
In Proceedings of the 1st Workshop on Evaluating NLG Evaluation, 2020
NUBIA is a modular learned metric for evaluating generated text. It combines neural feature extraction, aggregation, and calibration to achieve competitive machine translation evaluation and state-of-the-art image caption quality evaluation.
@inproceedings{kane2020nubia,title={NUBIA: NeUral Based Interchangeability Assessor for Text Generation},author={Kane, Hassan and Kocyigit, Muhammed Yusuf and Abdalla, Ali and Ajanoh, Pelkins and Coulibali, Mohamed},booktitle={Proceedings of the 1st Workshop on Evaluating NLG Evaluation},pages={28--37},year={2020},doi={10.18653/v1/2020.evalnlgeval-1.4},}
2019
Towards Neural Similarity Evaluator
Hassan Kane, Yusuf Kocyigit, Pelkins Ajanoh, Ali Abdalla, and Mohamed Coulibali
In NeurIPS 2019 Workshop on Document Intelligence, 2019
We examine limitations of BLEU and ROUGE, define behavioral criteria for strong similarity metrics, and show the potential of Transformer-based language models as learned evaluators for generated text.
@inproceedings{kane2019neural,title={Towards Neural Similarity Evaluator},author={Kane, Hassan and Kocyigit, Yusuf and Ajanoh, Pelkins and Abdalla, Ali and Coulibali, Mohamed},booktitle={NeurIPS 2019 Workshop on Document Intelligence},year={2019},}