🤖 AI Summary
This study addresses the puzzling observation that nonlinear models often fail to outperform linear counterparts on biomedical tabular data, despite the underlying biological mechanisms being highly nonlinear. The authors demonstrate that measurement noise preferentially degrades nonlinear signal structures, thereby masking the potential advantages of complex models. To formalize this insight, they establish the first systematic link between measurement error theory and the performance gap between linear and nonlinear models, proposing a tripartite framework governed by measurement reliability, sample size, and feature representation. They further derive an excess risk identity grounded in measurement error statistics and Gaussian analysis. Empirical validation across 140 UK Biobank prediction tasks confirms that the observed performance gap strongly reflects data noise characteristics, underscoring that neither scaling data volume nor increasing model complexity can compensate for information loss induced by low measurement reliability.
📝 Abstract
On biomedical tabular data, flexible models such as deep networks, gradient-boosted trees, and kernel methods are repeatedly matched or beaten by linear and logistic regression given the same features. The usual reaction is to treat this as a model-side shortfall, to be fixed with more data, a better architecture, or tuning, on the assumption that the nonlinear structure is there and the model has failed to capture it. We argue that these fixes cannot help when the binding limit is the measurement rather than the model, as it frequently is in biomedicine. Additive noise blurs the population-optimal predictor, and because blurring removes a function's fine, rapidly varying detail before its broad shape, it erases nonlinear structure faster than linear structure. A degree-$k$ interaction is attenuated by the $k$-th power of feature reliability, while the linear part is attenuated only once. At the reliabilities typical of biomedical measurement, the nonlinear advantage can vanish even when the underlying biology is strongly nonlinear, and what the noise removes cannot be recovered by a larger cohort or a more flexible model, only by better measurement. The nonlinearity is hidden, not absent, and a tie between linear and flexible models is not by itself a verdict on the biology. These pieces are classical, drawn from measurement-error statistics, psychometrics, and Gaussian analysis, and we assemble them into an exact excess-risk identity. Measurement reliability is one of three conditions, alongside sample size and feature representation, that must align for a flexible model to help, and together they leave only a narrow window that most biomedical tasks fall outside. Across 140 UK Biobank tasks, the gap between flexible and linear models, where it exists, carries the predicted noise signature, and the three conditions can be separated by intervention but not by a benchmark alone.