🤖 AI Summary
This study addresses the challenge of disentangling whether coefficient heterogeneity in multicenter prognostic models arises from differences in case-mix or site-specific contextual effects. The authors propose the first diagnostic framework capable of decoupling these sources: leveraging an autoencoder to construct a low-dimensional latent space that emphasizes local prognostic relationships, combined with a tailored loss function, site-specific local regressions, and coefficient surface decomposition. This approach separates heterogeneity into a cross-site reference and site-specific deviations, which are then mapped onto the outcome scale to generate both observation-level and site-level summaries. In a two-center COPD trial, the dominant source of slope heterogeneity in key latent variables was attributed to context, while observation-level variance was primarily driven by case-mix; site-level summaries highlighted contextual differences. Negative-control permutation tests confirmed the authenticity of the contextual signal, offering empirical guidance for choosing between unified or localized modeling strategies.
📝 Abstract
Prognostic regression models often synthesize data from multiple sites, whether within a multi-site study, across federated settings, or in individual participant data meta-analysis. Here, a site is any data source, such as a hospital, registry, trial, or study, and need not be a physical center. Analysts must then decide whether one regression model represents all sites or whether site-specific models are needed. Established measures such as coefficient-level tau^2 quantify heterogeneity but do not distinguish its source. We focus on diagnosing whether coefficient heterogeneity reflects case-mix or site-specific context effects. Case-mix heterogeneity can arise when linear regression terms approximate multivariable non-linear relationships in populations with different covariate distributions. Contextual heterogeneity arises when comparable patients require different regression relationships across sites. We do this by fitting site-specific local regressions in a dimension-reduced space and partitioning the smoothed coefficient surfaces into a cross-site reference and site-specific deviations. An autoencoder and custom loss structure the latent space around local prognostic relationships. We then project this partition onto the outcome scale to derive observation- and site-level summaries. We demonstrate the approach on a COPD trial with two sites. In the three leading latent slope coordinates, coefficient-surface variation was predominantly contextual. The derived observation-level outcome-scale variance partition was case-mix-leading, whereas its between-site aggregation was concentrated in contextual differences rather than case-mix shifts. A permuted-site negative control assesses whether the contextual summary can arise when site labels carry no signal. This diagnostic distinction can inform whether joint or site-specific regression models should be evaluated.