Machine Learning Seminar: Jeonghwan Kim, "Learning to See, Integrate, and Act: Visual Grounding for Multimodal Foundation Models."

Sep 25, 2026   2:00 pm  
1302 Siebel Center
Sponsor
Siebel School of Computing and Data Science
Speaker
Jeonghwan Kim
Contact
Weixin Chen
E-Mail
weixinc2@illinois.edu
Originating Calendar
Siebel School Speakers Calendar
Abstract: Multimodal foundation models have become increasingly capable at language-based reasoning, yet their visual representations often fail to preserve the dense, granular evidence needed to support that reasoning. Jeonghwan studies how to build multimodal models that remain tightly grounded in visual information while retaining access to the text-based knowledge that enables interpretation, reasoning, and decision-making. Across perception, retrieval, and action, he investigates how models can better identify relevant visual evidence, connect that evidence to internal and external knowledge, and use it to guide downstream reasoning in an embodied environment. More broadly, his work positions visual grounding as a core principle of multimodal intelligence: not simply localizing objects in an image, but maintaining a reliable connection between what a model sees, what it knows, and what it does.

Bio: Jeonghwan Kim is a Ph.D. candidate in Computer Science at the University of Illinois Urbana-Champaign, advised by Professor Heng Ji. His research focuses on multimodal foundation models that bridge fine-grained visual perception, reasoning, and embodied intelligence. Prior to his work on multimodal AI, he conducted research in natural language processing, including multi-hop question answering, retrieval-augmented generation, and numerical reasoning. His work has appeared in leading venues such as NeurIPS, ICLR, ACL, EMNLP, NAACL, and CVPR, including a NeurIPS 2025 Spotlight paper on part-level visual understanding in large multimodal models. He is also a 2026-2027 Capital One PhD Fellow and has interned with various companies such as Meta Reality Labs and Amazon.
link for robots only