single-cell differential expression modeling and pseudobulk methods

Develops statistical models and pseudobulk workflows for single-cell differential expression, designing analyses that aggregate single-cell measurements appropriately and produce interpretable DE results.

single-celldifferentialexpressionmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Aug 01, 2026Aug 01, 2026
Career
Value
No comparison yet
$200K/year
Aug 01, 2026Aug 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Running in circles: practical limitations for real-life application of data fission and data thinning in post-clustering differential analysis

May 22, 2024
BH
Benjamin Hivert
🏛️ Univ. Bordeaux | INSERM | INRIA | Vaccine Research Institute | Hôpital Henri Mondor | RAND Corporation | CHU Bordeaux | Service d'information médicale

In post-clustering differential expression analysis of scRNA-seq data, data splitting fails under unknown cluster structures: it relies on known cluster labels to estimate cluster-specific scale parameters, thereby violating independence and severely inflating Type I error. This work identifies its fundamental limitations—the non-mixture assumption is invalid under mixture distributions, and conditional data splitting requires prior cluster knowledge. To address this, we propose a neighborhood-based heteroscedastic nonparametric scale estimation framework. We provide the first theoretical quantification of how scale estimation bias propagates to Type I error inflation and derive a tight upper bound on the resulting error. We prove that standard data splitting inevitably fails without ground-truth clusters, whereas our method remains robust provided sufficient cluster separation. Our analysis rigorously establishes the practical applicability boundary of data splitting and delivers a verifiable, nonparametric alternative for post-clustering inference.

Addresses Type I error control in post-clustering differential expression analysisIdentifies violations of parametric assumptions in clustered scRNA-seq dataReveals fundamental limitations of data fission in real-world applications

Single-cell RNA-seq data are inherently noisy and sparse, and existing visualization methods often lack a solid statistical foundation, making it difficult to distinguish technical artifacts from genuine biological signals. This work proposes ZINBGT—a zero-inflated negative binomial mixture model with a geometric tail—that uniquely integrates statistical rigor with biological interpretability. The model provides interpretable visualizations of gene expression at the individual gene level and employs Wasserstein distance to diagnose aberrant genes. Applying ZINBGT, the authors successfully identify outlier-expressing genes in *T. brucei*, uncover intrinsic relationships among sparsity, mean expression, and dispersion in human immune cells, and highlight limitations of current simulated datasets in accurately recapitulating key characteristics of real single-cell data.

data noisesingle-cell transcriptomicsstatistical inference

Conventional single-cell RNA-seq (scRNA-seq) analyses relying solely on summary statistics (e.g., mean, variance) suffer from low sensitivity in detecting expression distribution differences and poor biological interpretability. Method: We propose a model-agnostic, tailored exponential family (SEF) density modeling framework that enables nonparametric estimation, visualization, and rigorous hypothesis testing of gene expression probability densities—both at the individual-cell and population levels—without assuming a specific parametric distribution. Leveraging asymptotically consistent covariance estimation and SEF-based statistical inference, our approach improves error control and statistical power. Contribution/Results: In comprehensive simulations across diverse scenarios, our method outperforms existing approaches. Applied to systemic lupus erythematosus (SLE) scRNA-seq data, it successfully identifies critical differentially expressed genes and functionally coherent gene sets missed by pseudo-bulk methods, establishing a novel paradigm for comparing single-cell expression distributions.

Detecting differential gene expression patterns missed by traditional bulk analysis methodsDeveloping framework for density estimation and group comparison without prior specificationsModeling individual gene expression as probability densities in single-cell genomics

This study addresses the lack of interpretable, auditable, and domain-informed automated reasoning methods in single-cell RNA sequencing analysis. The authors propose an “omics-native reasoning” paradigm and develop the first framework enabling large language models to directly invoke single-cell data and bioinformatics tools within natural language dialogues. Key tasks—such as cell type annotation, developmental trajectory reconstruction, and transcription factor target inference—are reformulated as iterative, stepwise reasoning processes that support correction and refinement. By integrating multi-turn reasoning, dynamic tool invocation, and evaluation on the scBench benchmark, the approach ensures transparent and traceable analytical logic. Experiments demonstrate that iterative reasoning improves cell type annotation accuracy by 11% over one-shot prompting, reduces graph edit distance by 30% in trajectory reconstruction using Gemini-2.5-Pro, and effectively resolves ambiguities in marker gene interpretation and regulatory mechanisms.

cell-type annotationdevelopmental trajectoryomics-native reasoning

Statistical Inference for Cell Type Deconvolution

Feb 13, 2022
DX
Dongyue Xie
🏛️ The University of Chicago

This study addresses the core challenges in cross-platform (bulk and scRNA-seq) cell type deconvolution: unreliable proportion estimation and absence of statistical inference. We propose the first theoretical framework ensuring identifiability of cell type proportions under arbitrary platform-specific biases. Methodologically, we develop a unified framework integrating mixed-effects modeling with asymptotic statistical inference, explicitly accounting for gene-wise correlations, technical scaling effects, measurement noise, and biological variability. Our approach enables asymptotically efficient inference for inter-individual proportion comparisons while rigorously controlling false discovery rates. Extensive simulations and analyses across multiple real-world bulk RNA-seq batches demonstrate substantial improvements in estimation accuracy and robustness. The method supports interpretable, statistically testable inference of cellular composition—filling a critical gap left by existing approaches, which neglect both inter-proportion dependencies and uncertainty quantification.

Addressing systematic scaling effects and data source differencesEstimating cell type composition across different measurement platformsProviding statistical inference for cell type proportion comparisons

Latest Papers

What's happening recently
View more

Modeling cellular responses to perturbations is highly challenging due to the heterogeneity of single-cell gene expression and complex implicit dependencies among genes. To address this, this work introduces, for the first time, a flow matching framework into the field, proposing an end-to-end method that directly models the effects of genetic and small-molecule perturbations on cell states within the native gene expression space. The approach employs a U-Net architecture to parameterize the velocity field and achieves high-fidelity predictions by fitting single-cell expression distributions. Evaluated on the PerturBench benchmark, the model demonstrates superior performance and was awarded first place in the general track of the inaugural ARC Virtual Cell Challenge, confirming its effectiveness and state-of-the-art capability in capturing both expression heterogeneity and perturbation-induced changes.

cell state modelingexpression heterogeneitygene dependencies

This work addresses the challenge of efficiently exploring and interpreting the high-dimensional combinatorial space of gene perturbations generated by AI-based virtual cell models and their complex transcriptional responses across diverse cell types. To this end, we propose a visual analytics system that, for the first time, integrates clustered overviews, compact glyph-based encodings, and coordinated multi-view interactions to enable systematic comparison and interpretable exploration of perturbation strategies. By incorporating AI-generated predictions and validating through real-world case studies and expert interviews, we demonstrate that our approach substantially enhances researchers’ understanding of perturbation effects and improves decision-making efficiency in drug discovery, effectively bridging the cognitive gap between computational models and biomedical experts.

drug discoverygene perturbationhigh-dimensional data

This study addresses the challenge of jointly modeling population-level variables (e.g., genotype) and observation-level variables (e.g., gene expression) in nested data structures such as individuals and their constituent cells. To this end, the authors propose the Nested Atomic Model (NAM), a Bayesian nonparametric approach that uniquely integrates both variable types within a unified hierarchical clustering framework, enabling simultaneous clustering of individuals and cells. An efficient variational Bayesian inference algorithm is developed to scale NAM to high-dimensional single-cell RNA sequencing and genotype datasets. Experiments on the OneK1K dataset demonstrate that NAM identifies groups of individuals with similar genotypes and consistent cell-type compositions, with cell-level clusters showing strong concordance with known immune cell types, thereby effectively uncovering multilayer biological heterogeneity.

group-level variableshierarchical datanested clustering

This work addresses the high storage and reuse costs of single-cell data and the lack of auditability in existing distillation methods that produce non-traceable synthetic data. The authors propose Minmax-CF, a method that, under fixed budgets of cells and genes, constructs a traceable core set by selecting only real observed cells through discrete minimax optimization and matching with static feature functions. This approach fully preserves original cell identifiers, gene symbols, and associated counts, labels, and metadata. Evaluated across multiple datasets, Minmax-CF achieves up to 96.52% balanced accuracy, substantially reduces pathway errors, and yields up to 2.55× GPU acceleration. The selected cells further enable efficient anomaly analysis and model validation.

auditabledata provenancereal-cell coresets

This work addresses the challenge that in single-cell perturbation data, cell populations under different perturbations exhibit substantial overlap, rendering conventional single-cell classification accuracy an unreliable metric of model performance. To overcome this limitation, the authors propose the Classifier Discrimination Score (CDS), which constructs a perturbation-level profile by aggregating classifier output probability distributions across entire cell populations and replaces single-cell predictions with population-level ranking. Remarkably, CDS recovers near-perfect perturbation identification from weak classifiers without requiring retraining. The method is compatible with diverse architectures—including linear models, MLPs, and Transformers—and demonstrates significant gains in identification accuracy on the Tahoe-100M and Virtual Cell Challenge datasets, with particularly pronounced advantages in low-cell-count regimes.

class overlapclassification accuracymodel evaluation

Hot Scholars

YZ

Yuanchun Zhou

Computer Network Information Center,CAS
Data MiningBig Data Analysis
ZZ

Zelin Zang

Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences
Deep Learning
YH

Yuankai Huo

Computer Science, Vanderbilt University
Medical Image AnalysisDeep LearningData Mining
RD

Ruining Deng

Weill Cornell Medicine
Medical Image AnalysisDeep LearningDigital Pathology