๐ค AI Summary
This paper addresses partial identification of target coefficients in linear regression when the outcome variable and a subset of covariates reside in two separately collected, non-linkable datasetsโwithout imposing exclusion restrictions. To overcome the limitation of conventional methods that rely on strong exogeneity assumptions, we first constructively characterize the sharp identification set under no exclusion constraints. We then propose a computationally efficient estimator for its bounds, based on moment inequalities and convex optimization. The estimator is analytically tractable, asymptotically normal, and exhibits robust finite-sample performance. Theoretically and empirically, our approach substantially extends the scope of prediction and causal inference in settings with missing individual-level linkage across data sources, offering a novel paradigm for modeling heterogeneous, multi-source data.
๐ Abstract
We study best linear predictions in a context where the outcome of interest and some of the covariates are observed in two different datasets that cannot be matched. Traditional approaches obtain point identification by relying, often implicitly, on exclusion restrictions. We show that without such restrictions, coefficients of interest can still be partially identified and we derive a constructive characterization of the sharp identified set. We then build on this characterization to develop computationally simple and asymptotically normal estimators of the corresponding bounds. We show that these estimators exhibit good finite sample performances.