publications

Research in language model evaluation, data contamination, representation learning, and computational social science.

2025

  1. Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination’s Impact on Machine Translation
    Muhammed Yusuf Kocyigit, Eleftheria Briakou, Daniel Deutsch, Jiaming Luo, Colin Cherry, and Markus Freitag
    In Proceedings of the 42nd International Conference on Machine Learning, 2025

2024

  1. Evaluation Data Contamination in LLMs: How Do We Measure It and (When) Does It Matter?
    Aaditya K. Singh, Muhammed Yusuf Kocyigit, Andrew Poulton, David Esiobu, Maria Lomeli, Gergely Szilvasy, and Dieuwke Hupkes
    arXiv preprint arXiv:2411.03923, 2024

2023

  1. A Novel Method for Analysing Racial Bias: Collection of Person Level References
    Muhammed Yusuf Kocyigit, Anietie Andy, and Derry Wijaya
    arXiv preprint arXiv:2310.15847, 2023
  2. Western, Religious or Spiritual: An Evaluation of Moral Justification in Large Language Models
    Eyup Engin Kucuk and Muhammed Yusuf Kocyigit
    In NeurIPS 2023 MP2 Workshop, 2023

2022

  1. AugCSE: Contrastive Sentence Embedding with Diverse Augmentations
    Zilu Tang, Muhammed Yusuf Kocyigit, and Derry Tanti Wijaya
    In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing, 2022
  2. Challenges in Measuring Bias via Open-Ended Language Generation
    Afra Feyza Akyurek, Muhammed Yusuf Kocyigit, Sejin Paik, and Derry Tanti Wijaya
    In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing, 2022
  3. On Measuring Social Biases in Prompt-Based Multi-Task Learning
    Afra Feyza Akyurek, Sejin Paik, Muhammed Yusuf Kocyigit, Seda Akbiyik, Serife Leman Runyun, and Derry Tanti Wijaya
    In Findings of the Association for Computational Linguistics: NAACL 2022, 2022
  4. Better Quality Estimation for Low Resource Corpus Mining
    Muhammed Kocyigit, Jiho Lee, and Derry Wijaya
    In Findings of the Association for Computational Linguistics: ACL 2022, 2022

2020

  1. NUBIA: NeUral Based Interchangeability Assessor for Text Generation
    Hassan Kane, Muhammed Yusuf Kocyigit, Ali Abdalla, Pelkins Ajanoh, and Mohamed Coulibali
    In Proceedings of the 1st Workshop on Evaluating NLG Evaluation, 2020

2019

  1. Towards Neural Similarity Evaluator
    Hassan Kane, Yusuf Kocyigit, Pelkins Ajanoh, Ali Abdalla, and Mohamed Coulibali
    In NeurIPS 2019 Workshop on Document Intelligence, 2019