Computer Vision Seminar Series: Dr. Huaizu Jiang, "Do 3D Vision Models Still Need Geometric Principles?"

Oct 9, 2026   12:30 pm  
2405 Siebel Center
Sponsor
Siebel School of Computing and Data Science
Originating Calendar
Siebel School Speakers Calendar
Abstract: Learned models such as VGGT now recover camera poses and 3D structure directly from images, a task that once required hand-designed geometric pipelines. Recent community discussions connect this progress to the Bitter Lesson, which holds that general methods eventually outperform methods built on human knowledge as computation grows. These discussions raise a fundamental question for 3D vision: do our models still need geometric principles, and if so, where? 

In this talk, I will present three projects ordered by how extensively each system relies on explicit geometry. Together, they examine how the value of geometric principles changes as the design priority shifts from generality to efficient inference. First, I will show that a network can relight a scene from several photos without any explicit scene geometry modeling. Explicit geometry appears only in the input, as rays that describe the camera viewpoints and the light sources. It thus serves as an interface that lets users set the lighting they want. Second, I will present a unified correspondence model that finds matchings across 2D–2D, 2D–3D, and 3D–3D pairs via one shared Transformer. It repurposes the attention matrix as matching cost and relies on the geometric structure of correspondences, pursuing more general vision models starting from one specific task. Third, I will describe a dense simultaneous localization and mapping (SLAM) system that uses geometric principles throughout, from stereo matching to camera tracking and loop closure. This design lets it run efficiently with a small memory footprint on an edge device, at some cost in accuracy. I will conclude with open questions about how geometric principles and large-scale learning can work together in future 3D vision models.

Speaker Bio.: Huaizu Jiang (https://jianghz.me/) is an assistant professor in the Khoury College of Computer Sciences at Northeastern University, where he leads the 3D Visual Intelligence Lab. His research centers on 3D computer vision, with connections to computer graphics and robotics. His long-term goal is to develop algorithmic foundations for machines to understand the geometry, dynamics, and semantics of the 3D world, so that agents can reason about and interact with complex physical environments. Before joining Northeastern, he was a postdoctoral researcher at Caltech and a visiting researcher at NVIDIA. He received his Ph.D. from the University of Massachusetts Amherst. His work was selected a Best Paper Award candidate at 3DV 2026. His research is supported by NSF and NIH.

Food will be served at 12PM before the talk.

There will also be a student roundtable session with Huaizu from 3-4 PM at Siebel 3102. Feel free to join!
link for robots only