Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分层审计方法,揭示心血管筛查模型的高准确性主要源于目标泄漏而非模型本身,并评估了透明模型在公平性和不确定性处理上的优势。
📝 Abstract
Cardiovascular screening models trained on national health surveys routinely report areas under the receiver operating characteristic curve (AUROC) near 0.89. We asked whether that accuracy reflects learning or target leakage, whether tabular foundation models change the answer, and whether the properties deployment requires survive joint examination. We benchmarked ten classifiers spanning linear, tree-ensemble, neural, glass-box, and tabular foundation classes for prevalent myocardial infarction in 442,067 respondents of the 2022 Behavioral Risk Factor Surveillance System across five feature tiers of decreasing leakage risk. Each was audited for discrimination, calibration, fairness at an explicit screening threshold, conformal coverage, explanation faithfulness, and inference cost, then applied -- models and thresholds frozen -- to 430,755 respondents of 2023. Removing two post-diagnostic features cost every model 0.049-0.051 AUROC, collapsing the field into a 0.0045-wide band. The glass-box explainable boosting machine was non-inferior to every alternative within a pre-specified 0.005 margin while scoring the cohort roughly 104 times faster than the strongest foundation model. One threshold detected 75.4% of women's infarctions against 89.0% of men's; editing the model's shape functions reduced the gap to 0.010. Marginal conformal prediction gave 0.86 coverage to men and 0.82 to adults over 60; Mondrian calibration repaired every stratum. Frozen models transported within 0.002 AUROC. Reported headroom in this literature is a property of the feature set, not the learner. Transparency cost nothing measurable and made fairness repair and uncertainty conditioning directly auditable. Evaluation practice, not model capacity, is the binding constraint.
Problem

Research questions and friction points this paper is trying to address.

target leakage
cardiovascular screening
foundation models
model accuracy
feature set
Innovation

Methods, ideas, or system contributions that make the work stand out.

target leakage
glass-box model
transparency
fairness
uncertainty conditioning
🔎 Similar Papers
R
Raad Bin Tareaf
Data Science and AI Cluster, XU Exponential University of Applied Sciences, Potsdam, Germany; German University of Digital Science, Potsdam, Germany
M
Murad Al-Rajab
College of Engineering, Abu Dhabi University, Abu Dhabi, United Arab Emirates
S
Samia Loucif
College of Technological Innovation, Zayed University, Abu Dhabi, United Arab Emirates
S
Samer Ellaham
Cleveland Clinic Hospital, Abu Dhabi, United Arab Emirates
C
Cedric Schmitz
Data Science and AI Cluster, XU Exponential University of Applied Sciences, Potsdam, Germany