Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures

- Sponsor
- Zhizhen Zhou, Ph.D.
- Speaker
- Maryam Fazel, Ph.D. - University of Washington
- Contact
- Zhizhen Zhou, Ph.D.
- zhizhenz@illinois.edu
Abstract:
Learning Gaussian Mixture Models (GMMs) is a fundamental problem in machine learning, and the Expectation-Maximization (EM) algorithm (Dempster,’77) and its variant gradient-EM are widely used algorithms for it. When the ground-truth GMM and the learning model have the same number of components m, a line of prior work has attempted to establish rigorous recovery guarantees; however, EM methods are known to fail to recover the ground truth when m>2.
This talk considers the “over-parameterized” case, where the learning model uses n>m components to fit an m-component GMM. I will show that gradient-EM converges globally and recovers the GMM: for a well-separated GMM with only mild over-parameterization n = \Omega(m log m), randomly initialized gradient-EM converges to the ground truth at a polynomial rate with polynomial samples. The analysis relies on novel characterization of the geometric landscape of the likelihood loss. This is the first global convergence result for EM methods beyond the special case of m=2. We will also discuss a way to speed up gradient-EM. More broadly, this talk highlights how over-parameterization or “scaling” can fundamentally alter optimization outcomes favorably for machine learning models.Bio:
Professor Fazel holds the Moorthy Family Inspiration Career Development Professorship in ECE, and adjunct appointments in the departments of Mathematics, Statistics, and the Allen School of Computer Science and Engineering at UW. Her current research interests are Optimization in Machine Learning and AI, Deep Learning Theory, Learning and Control, and Reinforcement learning. Professor Fazel is the director and lead PI of the Institute for Foundations of Data Science (IFDS), a multi-university research institute aiming to develop theoretical foundations for machine learning and data science, funded by a $12.5 million NSF TRIPODS Phase II grant, launched in September 2020. Previously, she co-directed the Algorithmic Foundations of Data Science Institute (ADSI), a TRIPODS Phase I institute that was a pre-cursor to IFDS. She is a recipient of the Farkas Prize of the INFORMS Optimization Society (2025), the NSF CAREER Award (2009), UWEE Outstanding Teaching Award (2009), and UAI conference Best Student Paper Award (with her student K. Dvijotham, 2014), and coauthored a paper on low-rank matrix estimation selected by ScienceWatch as the “Fast Breaking Paper” (based on citation numbers) in the area of Mathematics (August 2011). Prior to joining UW, she was a postdoctoral scholar at Caltech; and received her PhD in EE from Stanford University where she was advised by Prof. Stephen Boyd. She received her BS in EE from Sharif University of Technology in Iran. Professor Fazel was a Founding Associate Editor of the SIAM Journal on Mathematics of Data Science (SIMODS), a Program Chair of the ICML 2025 conference, and currently serves on the Editorial board of the MOS-SIAM Book Series on Optimization, the the Advisory board of the UW-Amazon Science Hub, and the executive committee of the eScience Institute. She is a co-organizer of the cross-campus seminar series Distinguished Seminars in Optimization and Data. She also works with Amazon as an Amazon Scholar.