Statistical Inference for Cell Type Deconvolution

📅 2022-02-13
📈 Citations: 6
Influential: 3
📄 PDF
🤖 AI Summary
This study addresses the core challenges in cross-platform (bulk and scRNA-seq) cell type deconvolution: unreliable proportion estimation and absence of statistical inference. We propose the first theoretical framework ensuring identifiability of cell type proportions under arbitrary platform-specific biases. Methodologically, we develop a unified framework integrating mixed-effects modeling with asymptotic statistical inference, explicitly accounting for gene-wise correlations, technical scaling effects, measurement noise, and biological variability. Our approach enables asymptotically efficient inference for inter-individual proportion comparisons while rigorously controlling false discovery rates. Extensive simulations and analyses across multiple real-world bulk RNA-seq batches demonstrate substantial improvements in estimation accuracy and robustness. The method supports interpretable, statistically testable inference of cellular composition—filling a critical gap left by existing approaches, which neglect both inter-proportion dependencies and uncertainty quantification.
📝 Abstract
Integrating data from different platforms, such as bulk and single-cell RNA sequencing, is crucial for improving the accuracy and interpretability of complex biological analyses like cell type deconvolution. However, this task is complicated by measurement and biological heterogeneity between target and reference datasets. For the problem of cell type deconvolution, existing methods often neglect the correlation and uncertainty in cell type proportion estimates, possibly leading to an additional concern of false positives in downstream comparisons across multiple individuals. We introduce MEAD, a comprehensive statistical framework that not only estimates cell type proportions but also provides asymptotically valid statistical inference on the estimates. One of our key contributions is the identifiability result, which rigorously establishes the conditions under which cell type proportions are identifiable despite arbitrary heterogeneity of measurement biases between platforms. MEAD also supports the comparison of cell type proportions across individuals after deconvolution, accounting for gene-gene correlations and biological variability. Through simulations and real-data analysis, MEAD demonstrates superior reliability for inferring cell type compositions in complex biological systems.
Problem

Research questions and friction points this paper is trying to address.

Estimating cell type composition across different measurement platforms
Addressing systematic scaling effects and data source differences
Providing statistical inference for cell type proportion comparisons
Innovation

Methods, ideas, or system contributions that make the work stand out.

MEAD framework enables accurate cell type proportion estimation
Identifiability result resolves platform-specific scaling differences
Supports cross-individual comparisons accounting for gene correlations
🔎 Similar Papers
No similar papers found.