Partial pooling predicts cross-validation reliability: a closed-form triage and Rao-Blackwellised cure for hierarchical LOO

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unreliability of Pareto-smoothed importance sampling leave-one-out cross-validation (PSIS-LOO) in hierarchical models when group-level random effects are data-driven, which can lead to misleading model assessments. To overcome this limitation, the authors propose a weight-free diagnostic for detecting fold failures, integrated with an observation-level Rao–Blackwellized leave-one-out (RB-LOO) estimator. By marginalizing over latent parameters via importance sampling and leveraging structural leverage combined with a closed-form triangular triage strategy, the method substantially enhances computational efficiency and numerical stability. Evaluated on logistic and count generalized linear mixed models, RB-LOO achieves threefold higher accuracy than moment matching and reproduces exact refitting results within 82 minutes (elpd RMSE = 0.04), thereby effectively preventing erroneous model selection.
📝 Abstract
For hierarchical models, Pareto-smoothed importance-sampling leave-one-out cross-validation (PSIS-LOO) fails on the folds where a random-effect coordinate is data-driven and its group is small. We show that the Gelman-Pardoe pooling factor and structural leverage predict these folds from model structure and group sizes, without forming importance weights. In Gaussian linear mixed models the leverage reduces to group size, giving a design-time map that separates the failing ($\hat{k}>0.7$) folds with AUC 0.96; across replicated logistic GLMMs the post-fit, weight-free predictor reaches AUC 0.81. The cure is integrated importance sampling: marginalise the random-effect block and importance-sample only the base parameters. This is not new, but we contribute its observation-level specialisation for random-intercept GLMMs: an analytic Gaussian downdate and a 1-D quadrature for Bernoulli, binomial and Poisson responses, packaged as a drop-in rb_loo(fit). Against exact refits, this marginalised estimator (RB-LOO) is $3\times$ more accurate than moment matching on singleton-heavy logistic GLMMs. On overdispersed count data with 97 failing folds, moment matching leaves 37 uncorrected and is no more accurate than raw PSIS-LOO, while RB-LOO reproduces the 82-minute exact refit (elpd RMSE 0.04) at no cost. The error changes decisions: against a negative-binomial model, PSIS-LOO reports decisive evidence ($z=4.9$) and reloo reports significant evidence ($z=3.4$) for the more complex model, where an exact analysis, reproduced by RB-LOO, finds the two indistinguishable ($z=1.0$). A base-fiber Schur decomposition splits case-deletion influence into a vertical (pooling) term that governs where PSIS-LOO fails and a horizontal (variance-component) term that governs where RB-LOO is itself strained, giving a two-level triage that recovers the exact answer while refitting only the few folds that need it.
Problem

Research questions and friction points this paper is trying to address.

hierarchical models
leave-one-out cross-validation
PSIS-LOO failure
small group sizes
random effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rao-Blackwellisation
hierarchical LOO
partial pooling
integrated importance sampling
structural leverage
💼 Related Jobs
No related jobs found.
A
Aidan D Bindoff
Wicking Dementia Research & Education Centre, University of Tasmania, Hobart, Tasmania, Australia