Learning Probabilities of Causation from Finite Population Data
Accurately estimating population-natural-stratum (PNS), probability of sufficiency (PS), and probability of necessity (PN) under limited observational or experimental data remains challenging due to reliance on full subgroup distribution assumptions. Method: This paper introduces the first machine learning framework for tight causal probability bound inference, integrating Tian–Pearl theoretical bounds with supervised learning over subgroup feature embeddings. Leveraging data from only ~500 observable subgroups, the method generalizes tightly bounded PNS estimates to 32,768 latent subgroups. Contribution/Results: The framework significantly reduces data requirements compared to conventional distribution-dependent approaches, enhancing feasibility of causal interpretability in small-sample settings. Its core innovation is establishing a novel “causal bound learning” paradigm—replacing the classical assumption of complete distributional knowledge with learnable, embedding-based generalization. This advances fine-grained causal assessment under finite-data constraints and opens new avenues for scalable, data-efficient causal inference.