Tailored for undergraduate researchers, this calendar is a curated list of research seminars at the University of Illinois. Explore the diverse world of research and expand your knowledge through engaging sessions designed to inspire and enlighten.

To have your events added or removed from this calendar, please contact OUR at ugresearch@illinois.edu

Machine Learning Seminar: Sagnik Mukherjee, "Optimization Geometry in LLM post-training."

Sep 4, 2026   2:00 - 3:15 pm  
1302 Siebel Center
Sponsor
Siebel School of Computing and Data Science
Speaker
Sagnik Mukherjee
Contact
Weixin Chen
E-Mail
weixinc2@illinois.edu
Originating Calendar
Siebel School Speakers Calendar

Abstract: On-policy training, including reinforcement learning (RL) and methods such as On-Policy Distillation (OPD/OPSD), has been a driving force behind the recent progress in LLM post-training. Despite their success in producing increasingly capable and general-purpose models, the distinctive optimization dynamics underlying these methods remain poorly understood. In this talk, Sagnik will argue that on-policy post-training exhibits unique properties in the nature and geometry of its weight updates. He will further present early explorations suggesting that commonly used optimizers may be poorly suited to these distinctive optimization dynamics. Finally, the talk will discuss how these observations can inform our understanding of modern post-training and potentially guide the development of improved post-training algorithms.

Bio: Sagnik is a third-year PhD student at UIUC, working with Prof. Dilek Hakkani-Tür and Prof. Hao Peng on post-training, reasoning, and sequential decision-making with large language models (LLMs). Prior to joining UIUC, he completed his undergraduate studies at IIT Kanpur. His earlier research has explored the dynamics and optimization of LLM post-training. In particular, his work has shown that (i) on-policy training induces sparse, structured updates to a base model (NeurIPS 2025), and (ii) reinforcement learning can benefit from simpler optimization methods, such as SGD without momentum; this work was selected for oral presentation at ICML 2026. More recently, Sagnik is interested in developing novel algorithms for long-horizon sequential decision-making with LLMs, particularly in settings involving partial observability and uncertainty. His research is guided by the view that insights from search and human cognition can provide valuable principles for designing more capable decision-making systems.

link for robots only