Learning Probabilities of Causation from Finite Population Data

📅 2022-10-16
🏛️ arXiv.org
📈 Citations: 7
Influential: 0
📄 PDF
🤖 AI Summary
Accurately estimating population-natural-stratum (PNS), probability of sufficiency (PS), and probability of necessity (PN) under limited observational or experimental data remains challenging due to reliance on full subgroup distribution assumptions. Method: This paper introduces the first machine learning framework for tight causal probability bound inference, integrating Tian–Pearl theoretical bounds with supervised learning over subgroup feature embeddings. Leveraging data from only ~500 observable subgroups, the method generalizes tightly bounded PNS estimates to 32,768 latent subgroups. Contribution/Results: The framework significantly reduces data requirements compared to conventional distribution-dependent approaches, enhancing feasibility of causal interpretability in small-sample settings. Its core innovation is establishing a novel “causal bound learning” paradigm—replacing the classical assumption of complete distributional knowledge with learnable, embedding-based generalization. This advances fine-grained causal assessment under finite-data constraints and opens new avenues for scalable, data-efficient causal inference.

Technology Category

Application Category

📝 Abstract
This paper deals with the problem of learning the probabilities of causation of subpopulations given finite population data. The tight bounds of three basic probabilities of causation, the probability of necessity and sufficiency (PNS), the probability of sufficiency (PS), and the probability of necessity (PN), were derived by Tian and Pearl. However, obtaining the bounds for each subpopulation requires experimental and observational distributions of each subpopulation, which is usually impractical to estimate given finite population data. We propose a machine learning model that helps to learn the bounds of the probabilities of causation for subpopulations given finite population data. We further show by a simulated study that the machine learning model is able to learn the bounds of PNS for 32768 subpopulations with only knowing roughly 500 of them from the finite population data.
Problem

Research questions and friction points this paper is trying to address.

Estimating causation probabilities with insufficient subpopulation data
Predicting PNS, PS, PN using machine learning models
Improving accuracy via MLP with Mish activation function
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine learning models predict causation probabilities
MLP with Mish activation reduces prediction error
Transfer learning from data-rich subpopulations
🔎 Similar Papers
No similar papers found.
A
Ang Li
Department of Computer Science, Florida State University
S
Song Jiang
Department of Computer Science, University of California Los Angeles
Yizhou Sun
Yizhou Sun
Professor, Computer Science, UCLA
Information NetworksKnowledge GraphsGraph Neural NetworksData MiningMachine Learning
J
J. Pearl
Department of Computer Science, University of California Los Angeles