Muhammed Yusuf Kocyigit

Muhammed Yusuf Kocyigit

Research Scientist at Google

Google

About

I am a research scientist at Google and a member of the Gemini large-scale pretraining core team. My work focuses on making language model evaluation more reliable: identifying and mitigating data contamination, improving evaluation signal-to-noise ratio, and understanding how training stages affect memorization and generalization. I also work on efficient, task-specific language models trained with synthetic data.

I completed my PhD in Computer Science at Boston University in 2025 under the supervision of Prof. Derry Wijaya. My doctoral research spanned LLM evaluation, machine translation, sentence representations, compositional generalization, and computational social science. I previously held research internships at Google, FAIR at Meta, and Amazon.

Interests
  • LLM Evaluation
  • Data Contamination and Decontamination
  • Model Generalization and Memorization
  • Pretraining Data Quality
  • Efficient Language Models
Education
  • PhD in Computer Science, 2020 - 2025

    Boston University

  • MSc in Electrical and Electronics Engineering, 2017 - 2020

    Bogazici University

  • BSc in Electrical and Electronics Engineering, 2012 - 2017

    Bogazici University

Experience

Research and industry

 
 
 
 
 
Research Scientist
2025 – Present Mountain View, CA
  • Analyze large-scale training runs for metric anomalies that may be caused by data contamination.
  • Develop improved decontamination recipes and test-train overlap metrics to identify contamination early and increase evaluation signal-to-noise ratio.
  • Study how training stages affect performance, memorization, and generalization.
  • Improve task-specific, efficient small language models using synthetic data.
 
 
 
 
 
PhD Research Fellow
2020 – 2025 Boston, MA
  • Studied natural language evaluation, with an emphasis on data contamination and its effect on reported model performance.
  • Improved sentence embeddings, compositional generalization evaluation, and cross-lingual representations through augmentation, multitask learning, and contrastive learning.
  • Conducted temporal bias research funded by the 2021 Google Research Scholar Program.
  • Led an Ottoman handwritten text recognition project that received $50K from the Scientific Research Council of Turkey.
 
 
 
 
 
Research Scientist Intern
2024 – 2025 Mountain View, CA
  • Quantified the effect of data contamination on LLM performance and evaluation reliability.
  • Built scalable experimental frameworks to measure contamination impact and improve evaluation data selection.
  • Host: Daniel Deutsch.
 
 
 
 
 
Research Scientist Intern
2023 – 2024 London, UK
  • Developed methods to define and detect contamination between evaluation sets and LLM pretraining data.
  • Analyzed performance overestimation caused by contamination; the methodology and code were used in contamination analysis for Llama 3.1.
  • Host: Dieuwke Hupkes.
 
 
 
 
 
Applied Science Intern
2022 – 2022 Cambridge, UK
  • Evaluated and improved compositional generalization in semantic parsing models.
  • Extended TMCD splits for more robust convergence and studied meta-learning and chain-of-thought prompting strategies.
  • Host: Emilio Monti.
 
 
 
 
 
Co-Founder and CEO
2018 – 2020 Istanbul, Turkey
  • Led business development, closed contracts worth more than $100K, and advised enterprise clients on AI strategy.
  • Deployed customer segmentation, lifetime value, and churn models that doubled targeted-ad conversion rates for LC Waikiki.
  • Improved illegal-electricity-use detection by 150% for DEDAS, with an estimated $500K business impact.

Publications

Research in language model evaluation, representation learning, and computational social science

(2025). Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation. Proceedings of the 42nd International Conference on Machine Learning (ICML 2025).

PDF Cite Source Document

(2024). Evaluation Data Contamination in LLMs: How Do We Measure It and (When) Does It Matter?. arXiv preprint.

PDF Cite Source Document

(2023). A Novel Method for Analysing Racial Bias: Collection of Person Level References. arXiv preprint.

PDF Cite Source Document

(2022). AugCSE: Contrastive Sentence Embedding with Diverse Augmentations. AACL-IJCNLP 2022.

PDF Cite Source Document DOI

(2022). Challenges in Measuring Bias via Open-Ended Language Generation. NAACL 2022, 4th Workshop on Gender Bias in Natural Language Processing.

PDF Cite Code Source Document DOI

(2022). On Measuring Social Biases in Prompt-Based Multi-Task Learning. Findings of NAACL 2022.

PDF Cite Source Document DOI

(2022). Better Quality Estimation for Low Resource Corpus Mining. Findings of ACL 2022.

PDF Cite Source Document DOI

(2020). NUBIA: NeUral Based Interchangeability Assessor for Text Generation. INLG 2020, Workshop on Evaluating NLG Evaluation.

PDF Cite Source Document DOI

(2019). Towards Neural Similarity Evaluator. NeurIPS 2019 Workshop on Document Intelligence.

Cite Source Document

Contact

The best way to reach me is by email.