Covariate Informed Identification of Heterogeneity and Outliers in Longitudinal Data

πŸ“… 2026-08-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Longitudinal data often exhibit multiple sources of heterogeneity, including divergent mean trajectories, increasing residual variance over time, and occasional outlying measurements. Conventional homogeneous models may yield inefficient parameter estimates and inflated variance assessments in such settings. This work proposes a novel Bayesian mixture model that, for the first time, incorporates covariate-driven binary indicator variables within a unified Bayesian framework to jointly model these three forms of heterogeneity via logistic regression. Inference is carried out using Markov chain Monte Carlo (MCMC) methods, and the approach facilitates posterior-probability-based model selection to evaluate the necessity of each heterogeneous component. Simulation studies demonstrate that the proposed method accurately identifies underlying heterogeneity structures and yields efficient fixed-effect estimates. Its practical utility is further corroborated through application to DHEAS hormone data from the Study of Women’s Health Across the Nation (SWAN).
πŸ“ Abstract
We often observe heterogeneity in longitudinal data, where the mean and variance for certain profiles meaningfully differ from the rest. Some profiles may also exhibit outliers at a limited number of measurements. Using a standard mixed effects model, which assumes homogeneity, can lead to overestimating the residual variance and inefficient estimation. In this work, we identify and account for three sources of heterogeneity in longitudinal data: incompatible mean trajectories, increased residual variance, and outliers at individual measurements. Our Bayesian mixture model incorporates binary indicators of heterogeneity for each of these features, modeled through logistic regression using covariates. We perform statistical inference using Markov chain Monte Carlo and implement model selection to evaluate the inclusion of various heterogeneous components. Simulations demonstrate that our model can accurately identify heterogeneity and produce efficient estimates of the fixed effects parameters. We further validate our approach using the DHEAS hormone data from the SWAN study.
Problem

Research questions and friction points this paper is trying to address.

heterogeneity
outliers
longitudinal data
mixed effects model
residual variance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayesian mixture model
heterogeneity identification
covariate-informed modeling
longitudinal data
outlier detection
πŸ”Ž Similar Papers
No similar papers found.
Anish Mukherjee
Anish Mukherjee
Assistant Professor, University of Liverpool
Theoretical Computer ScienceAlgorithmsComplexity Theory
J
Jeremy T. Gaskins
Department of Bioinformatics and Biostatistics, University of Louisville, Louisville, Kentucky, USA