🤖 AI Summary
Accurately estimating population-natural-stratum (PNS), probability of sufficiency (PS), and probability of necessity (PN) under limited observational or experimental data remains challenging due to reliance on full subgroup distribution assumptions. Method: This paper introduces the first machine learning framework for tight causal probability bound inference, integrating Tian–Pearl theoretical bounds with supervised learning over subgroup feature embeddings. Leveraging data from only ~500 observable subgroups, the method generalizes tightly bounded PNS estimates to 32,768 latent subgroups. Contribution/Results: The framework significantly reduces data requirements compared to conventional distribution-dependent approaches, enhancing feasibility of causal interpretability in small-sample settings. Its core innovation is establishing a novel “causal bound learning” paradigm—replacing the classical assumption of complete distributional knowledge with learnable, embedding-based generalization. This advances fine-grained causal assessment under finite-data constraints and opens new avenues for scalable, data-efficient causal inference.
📝 Abstract
This paper deals with the problem of learning the probabilities of causation of subpopulations given finite population data. The tight bounds of three basic probabilities of causation, the probability of necessity and sufficiency (PNS), the probability of sufficiency (PS), and the probability of necessity (PN), were derived by Tian and Pearl. However, obtaining the bounds for each subpopulation requires experimental and observational distributions of each subpopulation, which is usually impractical to estimate given finite population data. We propose a machine learning model that helps to learn the bounds of the probabilities of causation for subpopulations given finite population data. We further show by a simulated study that the machine learning model is able to learn the bounds of PNS for 32768 subpopulations with only knowing roughly 500 of them from the finite population data.