Statistics Seminar - Kaizheng Wang (Columbia) "Learning to augment statistical inference with generative models"

- Sponsor
- Department of Statistics
- Views
- 67
- Originating Calendar
- Department of Statistics Event Calendar
Title: Learning to augment statistical inference with generative models
Abstract: Generative models can produce seemingly unlimited synthetic data, yet discrepancies between synthetic and real populations can introduce bias and undermine statistical conclusions. For a new task with scarce or no real data, how much synthetic data can be safely used? This talk develops a general framework that leverages historical tasks to calibrate synthetic-data augmentation and quantify the resulting uncertainty. When no real observations are available, the framework adaptively selects the synthetic sample size to achieve nominal average-case coverage. In the application to LLM-generated survey responses, this calibrated size can be interpreted as the number of human respondents the LLM is effectively worth, providing a measure of its simulation fidelity. More generally, when real observations are available, the augmentation scheme is characterized jointly by the number of synthetic observations and the weight assigned to each. We learn a size–weight frontier that identifies the largest synthetic contribution compatible with reliable uncertainty quantification. These results show how generative models can strengthen statistical analysis when real data are limited.
