Chenyan Xiong

Associate Professor, Language Technologies Institute, Carnegie Mellon University.

prof_pic.jpg

6409 GHC,

5000 Forbes Avenue

Pittsburgh, PA 15213

I am an Associate Professor at the Language Technologies Institute (LTI) since 2023, with a courtesy/affiliate appointment in the Machine Learning Department (MLD), in the School of Computer Science at Carnegie Mellon University. I am also a co-founder and co-lead of the CMU Foundation and Language Model Center (FLAME). From 2018 to 2023, I worked at Microsoft Research Redmond on conversational search, dense retrieval, and large-scale pretraining, contributing scientific advances and real-world impact to production systems serving billions of users, trillions of web pages, billions of revenue boosts, and trillions of parameter frontier models. One of my main efforts back then was leading the scaling of MSFT’s first wave of large language models (the Megatron and DeepSpeed era), where our team built the cluster, developed large-scale training infrastructure, and converted much of our research into MSFT’s central LLMs shipped across various production systems. I also spent time part-time at Meta (2024), where we built large-scale foundation models for recommendation systems.

My research group welcomes Ph.D. students, postdoctoral researchers, and undergraduate/graduate interns. Recent publications are available at the CX Research Group at CMU. If our research interests align, please feel free to reach out.

  • Ph.D. students: I primarily review applicants in LTI’s Ph.D. program. You can list me as a potential advisor in the application system and send me an email to ensure I see your materials.
  • Postdocs: Please contact me directly via email.
  • Current CMU students: Fill out this form and email me. I particularly enjoy working with students who share my research interests, have well-defined directions, and value long-term impact or real-world applications.

Research Interests

My group’s current research aims to advance intelligence through foundation models, with specific efforts as follows:

Advance the core intelligence level of foundation models via:
  • Data strategies from pretraining data to posttraining environments, both organic and synthetic;
  • Self-improvement paradigms via auto-research loops spanning data, architecture, and multi-agent exploration;
  • Model–infrastructure co-design to advance the capability, efficiency, and scalability of foundation models.
Expand the intelligence to other frontiers, mainly healthcare, physical intelligence, and sports, through:
  • Neural architectures necessary to model new modalities from other frontiers;
  • Pretraining paradigms that achieve the unsupervised-learning advantages seen in language;
  • Posttraining approaches to enhance the capabilities of foundation models in specific scenarios.
Understand the implications of rapidly growing intelligence capabilities in the world, such as:
  • New business models and economic designs for AI-native applications;
  • Impact on the digital economy, the future of work, and the fairness of economic distribution in the new value chain;
  • New scientific research paradigms with the rise of AI.