Score
Techniques to detect, quantify, and mitigate collinearity among predictors—using tools like SEM‑derived construct scores, cluster‑level models, and diagnostic benchmarks—to produce transparent OLS comparisons and robust causal estimates. This includes specifying Bayesian marketing‑mix models and assessing whether clustering or reparameterization reduces multicollinearity.
This study addresses the challenge of identifying independent causal effects in observational causal inference when explanatory variables exhibit high multicollinearity. To mitigate this issue, the authors propose a novel approach that integrates hierarchical clustering with a Bayesian marketing mix model. The method first groups geographic units into clusters based on the correlation structure of marketing expenditures, thereby constructing cluster-level data that eliminate shared temporal trends. Causal effects of individual marketing channels are then estimated at the cluster level. This work represents the first application of hierarchical clustering specifically designed to alleviate multicollinearity in causal inference, achieving substantially improved identification accuracy while preserving the interpretability of causal estimates. Empirical results demonstrate that the proposed framework effectively disentangles and accurately quantifies the distinct causal impacts of different advertising channels.
In multi-treatment, partially overlapping observational causal inference settings—such as multi-channel marketing or multi-category ordering—conventional approaches that estimate treatment effects independently suffer from high variance and weak decision support. This paper proposes a customized ridge regression framework that employs a data-driven shrinkage strategy to adaptively balance heterogeneity and homogeneity across treatment effects. The method reduces the mean squared error of individual effect estimates while preserving interpretability and reconstructibility of aggregated causal effects. Theoretical analysis establishes consistency and convergence of the estimator; simulation studies corroborate its finite-sample performance. Real-world deployment at Wayfair demonstrates substantial improvements in causal decision accuracy and operational reliability for multi-treatment scenarios.
Conventional clustered robust inference fails when cluster sizes are non-negligible—e.g., following Zipf’s law—and 77% of empirical studies in the *American Economic Review* and *Econometrica* (2020–2021) violate its implicit equal-size or bounded-size assumptions. Method: This paper establishes the first necessary and sufficient condition for consistency of clustered robust estimators and proposes two new procedures: score subsampling and size-adjusted reweighting. Both methods are theoretically grounded—guaranteeing consistency and uniform size control—and practically implementable, with ready-to-use Stata packages. Results: Monte Carlo simulations demonstrate that the proposed methods strictly maintain nominal test size even where conventional approaches severely distort inference. They constitute the first truly robust and implementable inferential framework for settings with large, heterogeneous cluster sizes.
In high-dimensional covariate modeling, parameter rank deficiency arising from multicollinearity undermines identification, and conventional sparsity or discrete heterogeneity assumptions often violate economic theory, leading to severe estimation bias. This paper proposes a novel joint estimation framework that sequentially learns high-dimensional parameters and an adaptive projection matrix—marking the first method to map unidentifiable high-dimensional parameters into a low-dimensional space amenable to consistent estimation. The approach preserves original parameter accuracy under rank-deficient conditions and enjoys rigorous consistency and asymptotic normality guarantees. Validated via sequential algorithms, high-dimensional projection learning, and Monte Carlo simulations, the method substantially reduces bias and improves estimation precision. Empirically, it reveals positive R&D spillover effects among firms, though private returns remain dominant.
This study addresses the inconsistency in causal effect estimates between observational studies and randomized controlled trials (RCTs) by proposing the first unified framework for decomposing causal effect heterogeneity. The framework systematically identifies and quantifies three sources of heterogeneity: differences in covariate distributions, variation in mediating pathways, and shifts in outcome-generating mechanisms. Methodologically, it formally defines effect decomposition across data types (observational vs. experimental), integrating causal inference, sensitivity analysis, and decomposition modeling, while enabling robust parameter estimation under multiple hypotheses. Evaluated through simulation studies and an empirical analysis of the “Moving to Opportunity” experiment, the framework demonstrates improved interpretability, robustness, and policy generalizability in synthesizing evidence from heterogeneous data sources.
This study addresses the bias in regression coefficient inference that arises when within-cluster dependence is ignored in clustered data. The authors propose a novel estimator that explicitly models the intra-cluster dependence structure, accommodating both fixed and diverging cluster sizes. The framework is further extended to random-coefficient models to enable inference on average effects and testing of linear hypotheses. Key contributions include a robust estimation approach that accounts for within-cluster correlation, a new Wald-type test designed for improved stability in high-level hierarchical parameters, and a theoretical demonstration of the inconsistency of pooled ordinary least squares (POLS) under random-coefficient settings. The validity and practical utility of the proposed methodology are substantiated through rigorous theoretical analysis, simulation studies, and empirical applications.
This study addresses the challenge that existing statistical methods struggle to effectively estimate average treatment effects in experiments involving both randomly assigned treatments and fixed covariates. The authors develop a unified theoretical framework that, for the first time, establishes a general estimating equation theory for misspecified linear regression models with mixed regressors—combining random treatment indicators and fixed covariates—and extends this framework to clustered data settings. By integrating estimating equations, misspecification-robust analysis, and causal inference techniques, the proposed approach yields valid causal interpretations of regression coefficients and their standard errors even under model misspecification. This methodology is broadly applicable to practical experimental designs, including completely randomized trials.
This study addresses the sensitivity of causal effect estimation to model misspecification in longitudinal cluster-randomized and quasi-experimental designs. Within an M-estimation framework, it demonstrates that fixed-effects models yield consistent and asymptotically normal estimates of nonparametrically defined treatment effects, provided the treatment effect structure is correctly specified—even when other model components are arbitrarily misspecified. The work establishes, for the first time, that fixed-effects models are valid for estimating superpopulation marginal effects and reveals their robustness to partial misspecification of the treatment effect structure across diverse longitudinal settings. Through theoretical analysis, simulations, and reanalyses of empirical data, the paper further shows that fixed-effects models outperform mixed-effects models in robustness and reliability when time-invariant confounding exists at the cluster or individual level.
This study addresses the challenge of identifying and estimating causal effects under network interference, where an individual’s treatment may spill over and affect others’ outcomes. The authors propose a solution based on a linear outcome model that yields unbiased and consistent estimates of both binary and continuous treatment effects when the interference structure is known or partially known. The approach accommodates both fixed and random interference network specifications and innovatively eliminates interference-induced bias while remaining compatible with standard linear regression software. It also conveniently allows for the incorporation of random effects and heteroskedasticity- and autocorrelation-consistent (HAC) standard errors. Numerical simulations and empirical analyses demonstrate the method’s effectiveness in bias correction and practical applicability.
This study addresses the bias in inference on quadratic forms of linear regression coefficients under high-dimensional covariates in clustered data—arising in contexts such as instrumental variable regression, variance component estimation, and testing multiple constraints. The authors propose an unbiased estimator based on leave-one-cluster-out (LOCO) cross-fitting and establish its asymptotic normality. They further develop novel leave-two- and leave-three-cluster variance estimators that ensure conservativeness under weaker conditions while maintaining computational efficiency. The proposed framework accommodates growing cluster sizes and covariate dimensions with sample size, making it suitable for settings with strong within-cluster dependence and high-dimensional covariates. Both theoretically and computationally, the method outperforms conventional plug-in estimators, delivering robust, consistent, and efficient cluster-robust inference.