Score
Designing and running targeted experiments or interventions that isolate causal effects by controlling confounds and measuring outcomes, and validating whether internal model or learned representations encode mechanistic structure beyond surface inputs. This covers field-study design, perturbation tests, and validation strategies to attribute observed effects to the intended manipulations.
When randomized controlled trials are infeasible—as in rare diseases or oncology—effectively leveraging external data becomes a critical challenge. This work proposes a six-step causal inference–based scientific framework that systematically integrates external control arm designs from single-arm and hybrid control trials, unifying Bayesian dynamic borrowing, frequentist approaches, and modeling of covariate shift and outcome drift. Centered on causal identifiability, the framework clarifies, for the first time, the trade-off between efficiency and robustness inherent in external control methodologies and underscores the necessity of sensitivity analyses in regulatory decision-making. Through a systematic literature review and empirical evaluation, the study provides a coherent guide and accompanying software tools to support both theoretical integration and practical application of external data.
To address latent bias arising from unmeasured confounding in observational studies, this paper proposes a novel causal inference paradigm based on sample splitting: data are partitioned into planning and analysis samples, where the former adaptively selects robust design parameters (e.g., matching strategies, covariate sets), and the latter yields unbiased causal estimates. The method innovatively integrates multiple testing correction, heteroskedasticity-robust covariance estimation, and formal sensitivity analysis—extending support to multiple outcome variables, thereby relaxing the conventional single-outcome assumption. We establish theoretical guarantees of statistical validity under latent bias. Simulation studies demonstrate substantially higher statistical power than benchmark methods under strong unmeasured confounding. Empirical application to assessing the multidimensional impacts of floods in Bangladesh confirms practical feasibility and robustness.
In high-dimensional observational settings of randomized controlled trials (RCTs), causal treatment effect estimation is prone to modeling and sampling selection biases. Method: We introduce ISTAnt—the first real-world visual causal benchmark grounded in ant-behavior RCTs—and theoretically prove that classification accuracy cannot serve as a proxy for causal estimation quality. We further propose representation-learning principles tailored for scientific causal inference. Contribution/Results: Through tripartite validation—rigorous theoretical analysis, controlled synthetic experiments, and real biological experiments—across 6,480 large-scale fine-tuned models built upon state-of-the-art vision backbones, we demonstrate that common deep learning practices (e.g., loss design, data sampling) induce substantial systematic biases, and classification performance exhibits no strong correlation with causal estimation accuracy. Our work establishes a reproducible benchmark, theoretical criteria, and practical guidelines for high-dimensional causal inference.
This study addresses the challenge of integrating randomized or single-arm clinical trials with external experimental or observational data to enable cross-study treatment comparisons and improve estimation precision of treatment effects. Methodologically, building upon the potential outcomes framework, we first develop a unified identification strategy for hybrid-data designs, systematically characterizing identifiability conditions across diverse designs—including historical controls, synthetic controls, and anchoring estimators—and propose a generalizable taxonomy of such designs along with corresponding causal inference principles. Our contribution lies in filling a critical theoretical gap in regulatory science regarding the rigorous integration of external controls, thereby establishing a methodological foundation for leveraging real-world evidence to complement trial-based evidence in pharmaceutical and medical device evaluation. This advancement significantly enhances the transportability of evidence and its applicability to regulatory decision-making.
This paper addresses the challenge of confounder selection in observational studies by proposing an interactive, iterative method that requires neither a pre-specified causal graph nor a complete set of candidate variables. Grounded in latent projection theory, the method dynamically expands a causal graph through successive user-provided local adjustment sets and automatically identifies a minimal “principal adjustment set,” thereby determining whether confounding is controllable. Its key contributions are threefold: (1) it is the first approach to achieve sound and complete confounding control assessment without prior structural assumptions on the causal graph; (2) it makes no assumptions about causal relationships among potential confounders; and (3) it bridges theoretical rigor with practical feasibility. Both theoretical analysis and empirical evaluation demonstrate that, under correct user feedback, the algorithm accurately identifies admissible adjustment sets and correctly determines confounding controllability.
This study addresses the problem of extrapolating causal effects from multi-site randomized controlled trials (RCTs) to a new target site with baseline survey data only. To handle site-level population heterogeneity and unobserved confounding, we propose modeling baseline covariates as functional data—thereby capturing site-specific confounding structures—for the first time. We then develop a design-oriented, nonparametric method to construct an optimal finite-dimensional feature space, ensuring optimal convergence rates for conditional average treatment effect (CATE) estimation. Our approach integrates functional data analysis, nonparametric regression, and causal transfer learning theory. Evaluated across five integrated multi-site RCTs on cash transfer programs, the method significantly improves prediction accuracy of treatment effects at target sites and quantifies the estimation gain attributable to adaptive transfer.
This study addresses the limitations of traditional control-based causal inference methods—such as matching and difference-in-differences—in settings characterized by pervasive or structurally ambiguous spillover effects, where reliance on uncontaminated control units impedes accurate identification of both average direct and spillover effects. Within the potential outcomes framework, this work provides the first systematic comparison between control-based and prediction-based counterfactual approaches—including interrupted time series and machine learning control—in terms of their identification capabilities. Through simulation and empirical analyses, the authors demonstrate that in environments with widespread interference, prediction-based methods can more reliably estimate certain causal parameters over short horizons, circumventing the stringent assumption of unperturbed units and thereby offering a promising alternative for causal inference under complex interference.
This study addresses the challenge of identifying heterogeneous treatment effects while controlling the false discovery rate (FDR) in matched observational studies, where limited sample sizes and unmeasured confounding pose significant obstacles. The authors propose a novel method that adaptively discovers interpretable subgroups defined by covariate thresholds under many-to-one matching designs. Their approach achieves exact FDR control at the subgroup level for the first time and integrates sensitivity analysis models to account for unobserved confounding, leveraging multiple controls to enhance statistical power. Theoretical analysis, simulations, and empirical evaluation demonstrate that the method outperforms existing baselines in both accuracy and power when estimating heterogeneous economic returns to college education.
This study addresses the problem of testing whether a treatment effect operates entirely through observed mediators and identifying causal mechanisms under control for covariates. The authors propose a statistical test based on double machine learning, extending— for the first time—the joint evaluation of full mediation and causal mechanism identification to non-randomized treatment settings. By integrating conditional independence testing, the method achieves root-n consistent and asymptotically normal inference even in the presence of high-dimensional covariates. Simulation studies demonstrate favorable finite-sample performance, and the approach is successfully applied to two randomized experiments examining maternal mental health and social norms.
This study addresses the failure of conventional causal inference methods in group interaction experiments, where within-group interactions and interference effects violate standard assumptions. The authors develop a design-based causal inference framework that systematically characterizes identifiability under various scenarios—such as fixed or random group assignment and presence or absence of interference—and proposes corresponding inference strategies. Innovatively, they introduce a coupling strategy to handle complex dependence structures, integrating sparse-sampling asymptotics, cluster-robust inference, and the potential outcomes framework. They demonstrate that, even under interference, cluster-robust methods consistently estimate marginalized exposure effects. Moreover, when interference is absent and assignment is randomized, the framework naturally reduces to the standard individual-level randomized experiment, thereby preserving compatibility with classical individual-level inference.
This study addresses the challenge of identifying direct, indirect, and total effects in spatial interference experiments, where spillover effects—particularly under cluster randomization—render the level of the spillover kernel non-identifiable. The authors develop a unified causal framework that expresses estimands as linear functionals of the spillover kernel and introduce an “anchoring assumption” to resolve the level identification problem. By positing separability between the kernel’s shape and its level, they demonstrate that all estimators inherently depend on the choice of anchor and propose a “level leverage” measure to quantify sensitivity to the unidentifiable level. Integrating linear exposure mappings, bias decomposition, and design augmentation strategies, the framework leverages prior geometric structure to compute leakage terms and shape errors. The work further shows that conventional cluster-based analyses are special cases of implicit anchoring, thereby unifying existing approaches and offering principled criteria for deciding whether to augment the experimental design or select an anchor.