A Hilbert-Valued Functional Decomposition Framework for Explaining Time-Dependent Outputs
本文针对时间依赖性输出的解释问题,提出了一种基于希尔伯特值函数分解框架的方法,能够提供考虑时间依赖性的多粒度解释。
本文针对时间依赖性输出的解释问题,提出了一种基于希尔伯特值函数分解框架的方法,能够提供考虑时间依赖性的多粒度解释。
This study addresses the trade-off between interpretability and flexibility in hybrid models, where neural networks often compromise model transparency. We propose a Generalized Linear Model augmentation framework based on Lipschitz-constrained invertible residual networks. By incorporating a controllable bias mechanism and posterior orthogonalization, the method achieves flexible nonlinear estimation and distribution correction while strictly preserving stochastic monotonicity and model identifiability. This approach establishes a semi-structured hybrid modeling paradigm that integrates high interpretability with flexibility, enabling user-defined trade-offs and quantifiable model bias. Consequently, it effectively mitigates the transparency bottleneck inherent in complex data modeling tasks.
This study addresses the challenge of predicting postoperative discharge destinations in older adults, a task hindered by limited clinical data that impedes perioperative care optimization. For the first time in geriatric surgery, the authors employ Adversarial Random Forests (ARF) to generate synthetic data, which is then integrated with real-world observations using multiple imputation strategies. The combined dataset is used to train logistic regression, random forest, and TabPFN models. Results demonstrate that data augmentation substantially enhances the performance of simpler models: logistic regression accuracy improves from 0.70 to 0.81, and AUC rises from 0.85 to 0.92. In contrast, more complex models—random forest and TabPFN—maintain consistently high performance, achieving approximately 0.84 accuracy and 0.94 AUC. This work validates the efficacy and practical utility of generative machine learning approaches for clinical prediction tasks under data scarcity.
This study addresses the challenge of automatically extracting interpretable prognostic morphological features from whole-slide images of IDH wild-type glioblastoma to predict patient survival. We propose a novel interpretable multiple instance learning framework that, for the first time, integrates a sparse autoencoder with a Cox proportional hazards model. Evaluated on 720 real-world cases, the model achieves an AUC of 0.67 in survival stratification. It identifies 24 visual patches significantly associated with survival, 21 of which were validated by neuropathologists and categorized into seven distinct histological feature classes. This approach enables an automatic mapping from raw histopathology images to clinically comprehensible morphological patterns, substantially enhancing model transparency and pathological relevance.
Missing values are pervasive in real-world data, yet existing interpretable machine learning (IML) methods commonly rely on single imputation, neglecting how imputation-induced uncertainty affects explanation stability—particularly confidence interval coverage probabilities. This work systematically quantifies the impact of imputation strategies on the coverage of confidence intervals for three canonical IML methods: permutation importance, partial dependence plots, and Shapley values. We compare single imputation (mean, median, regression) against multiple imputation (MICE), integrating permutation testing and nonparametric confidence interval estimation. Results show that single imputation severely underestimates variance, yielding actual 95% confidence interval coverage rates frequently below 80%. In contrast, multiple imputation restores coverage close to the nominal level across most settings, substantially enhancing the statistical reliability of IML explanations. To our knowledge, this is the first study to rigorously evaluate and quantify how imputation choices affect the inferential validity of IML outputs.
本文针对时间依赖性输出的解释问题,提出了一种基于希尔伯特值函数分解框架的方法,能够提供考虑时间依赖性的多粒度解释。
This study addresses the trade-off between interpretability and flexibility in hybrid models, where neural networks often compromise model transparency. We propose a Generalized Linear Model augmentation framework based on Lipschitz-constrained invertible residual networks. By incorporating a controllable bias mechanism and posterior orthogonalization, the method achieves flexible nonlinear estimation and distribution correction while strictly preserving stochastic monotonicity and model identifiability. This approach establishes a semi-structured hybrid modeling paradigm that integrates high interpretability with flexibility, enabling user-defined trade-offs and quantifiable model bias. Consequently, it effectively mitigates the transparency bottleneck inherent in complex data modeling tasks.
This study addresses the challenge of predicting postoperative discharge destinations in older adults, a task hindered by limited clinical data that impedes perioperative care optimization. For the first time in geriatric surgery, the authors employ Adversarial Random Forests (ARF) to generate synthetic data, which is then integrated with real-world observations using multiple imputation strategies. The combined dataset is used to train logistic regression, random forest, and TabPFN models. Results demonstrate that data augmentation substantially enhances the performance of simpler models: logistic regression accuracy improves from 0.70 to 0.81, and AUC rises from 0.85 to 0.92. In contrast, more complex models—random forest and TabPFN—maintain consistently high performance, achieving approximately 0.84 accuracy and 0.94 AUC. This work validates the efficacy and practical utility of generative machine learning approaches for clinical prediction tasks under data scarcity.
This study addresses the challenge of automatically extracting interpretable prognostic morphological features from whole-slide images of IDH wild-type glioblastoma to predict patient survival. We propose a novel interpretable multiple instance learning framework that, for the first time, integrates a sparse autoencoder with a Cox proportional hazards model. Evaluated on 720 real-world cases, the model achieves an AUC of 0.67 in survival stratification. It identifies 24 visual patches significantly associated with survival, 21 of which were validated by neuropathologists and categorized into seven distinct histological feature classes. This approach enables an automatic mapping from raw histopathology images to clinically comprehensible morphological patterns, substantially enhancing model transparency and pathological relevance.
Missing values are pervasive in real-world data, yet existing interpretable machine learning (IML) methods commonly rely on single imputation, neglecting how imputation-induced uncertainty affects explanation stability—particularly confidence interval coverage probabilities. This work systematically quantifies the impact of imputation strategies on the coverage of confidence intervals for three canonical IML methods: permutation importance, partial dependence plots, and Shapley values. We compare single imputation (mean, median, regression) against multiple imputation (MICE), integrating permutation testing and nonparametric confidence interval estimation. Results show that single imputation severely underestimates variance, yielding actual 95% confidence interval coverage rates frequently below 80%. In contrast, multiple imputation restores coverage close to the nominal level across most settings, substantially enhancing the statistical reliability of IML explanations. To our knowledge, this is the first study to rigorously evaluate and quantify how imputation choices affect the inferential validity of IML outputs.