Design-based Inference for Treatment Effects in Observational Studies using Estimated Propensity Scores

- Sponsor
- Econometrics
- Speaker
- Panagiotis (Panos) Toulis (Chicago Booth)
- Views
- 1
- Originating Calendar
- Econometrics (SEMINARS)
Abstract: We study randomization inference for treatment effects in observational data under unconfoundedness. The main testing procedure is design-based in that it holds outcomes and covariates fixed and repeatedly re-randomizes treatment assignment according to estimated propensity scores. Under the sharp null hypothesis of no treatment effect, we derive non-asymptotic bounds on excess Type I error for tests based on an arbitrary test statistic in terms of in-sample and out-of-sample estimation error. Under the weak null hypothesis of no average treatment effect, we show that our tests are asymptotically valid when based on suitable studentized estimators of the average treatment effect. In contrast to earlier work on studentization in related settings, a key challenge in our analysis is adequate control of the error stemming from the estimation of the propensity score for each individual observation. For this reason, naive studentization may fail to be asymptotically valid for the weak null hypothesis, but we show that validity can be achieved by studentizing a score-adjusted inverse propensity-weighted estimator of the average treatment effect in the parametric case, or a Neyman-orthogonal estimator in the nonparametric case. We then compare the randomization test with the corresponding test using a standard normal critical value. Under the weak null hypothesis, we find that they are first-order equivalent and that neither test dominates the other in a higher-order comparison based on Edgeworth expansions. The randomization test is, however, more accurate to higher-order in regimes "near the sharp null,'' i.e., in which treatment effects are sufficiently small or sparse. While randomization inference is common in experimental settings, our results support its use for testing weak null hypotheses in observational studies as well, especially when treatment effects are believed to lie in the near-sharp regime.
