Institution profile

Leibniz Institute for Prevention Research and Epidemiology – BIPS

Academic institutioneurope · de
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models

Aug 17, 2026

This study addresses the trade-off between interpretability and flexibility in hybrid models, where neural networks often compromise model transparency. We propose a Generalized Linear Model augmentation framework based on Lipschitz-constrained invertible residual networks. By incorporating a controllable bias mechanism and posterior orthogonalization, the method achieves flexible nonlinear estimation and distribution correction while strictly preserving stochastic monotonicity and model identifiability. This approach establishes a semi-structured hybrid modeling paradigm that integrates high interpretability with flexibility, enabling user-defined trade-offs and quantifiable model bias. Consequently, it effectively mitigates the transparency bottleneck inherent in complex data modeling tasks.

0 citationsRead paper

Improving post-operative discharge destination prediction of geriatric patients with generative data augmentation

Apr 19, 2026

This study addresses the challenge of predicting postoperative discharge destinations in older adults, a task hindered by limited clinical data that impedes perioperative care optimization. For the first time in geriatric surgery, the authors employ Adversarial Random Forests (ARF) to generate synthetic data, which is then integrated with real-world observations using multiple imputation strategies. The combined dataset is used to train logistic regression, random forest, and TabPFN models. Results demonstrate that data augmentation substantially enhances the performance of simpler models: logistic regression accuracy improves from 0.70 to 0.81, and AUC rises from 0.85 to 0.92. In contrast, more complex models—random forest and TabPFN—maintain consistently high performance, achieving approximately 0.84 accuracy and 0.94 AUC. This work validates the efficacy and practical utility of generative machine learning approaches for clinical prediction tasks under data scarcity.

0 citationsRead paper

Explainable histomorphology-based survival prediction of glioblastoma, IDH-wildtype

Jan 16, 2026

This study addresses the challenge of automatically extracting interpretable prognostic morphological features from whole-slide images of IDH wild-type glioblastoma to predict patient survival. We propose a novel interpretable multiple instance learning framework that, for the first time, integrates a sparse autoencoder with a Cox proportional hazards model. Evaluated on 720 real-world cases, the model achieves an AUC of 0.67 in survival stratification. It identifies 24 visual patches significantly associated with survival, 21 of which were validated by neuropathologists and categorized into seven distinct histological feature classes. This approach enables an automatic mapping from raw histopathology images to clinically comprehensible morphological patterns, substantially enhancing model transparency and pathological relevance.

0 citationsRead paper

Imputation Uncertainty in Interpretable Machine Learning Methods

Dec 19, 2025

Missing values are pervasive in real-world data, yet existing interpretable machine learning (IML) methods commonly rely on single imputation, neglecting how imputation-induced uncertainty affects explanation stability—particularly confidence interval coverage probabilities. This work systematically quantifies the impact of imputation strategies on the coverage of confidence intervals for three canonical IML methods: permutation importance, partial dependence plots, and Shapley values. We compare single imputation (mean, median, regression) against multiple imputation (MICE), integrating permutation testing and nonparametric confidence interval estimation. Results show that single imputation severely underestimates variance, yielding actual 95% confidence interval coverage rates frequently below 80%. In contrast, multiple imputation restores coverage close to the nominal level across most settings, substantially enhancing the statistical reliability of IML explanations. To our knowledge, this is the first study to rigorously evaluate and quantify how imputation choices affect the inferential validity of IML outputs.

0 citationsRead paper
Recent publications

Latest Papers

LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models

Aug 17, 2026

This study addresses the trade-off between interpretability and flexibility in hybrid models, where neural networks often compromise model transparency. We propose a Generalized Linear Model augmentation framework based on Lipschitz-constrained invertible residual networks. By incorporating a controllable bias mechanism and posterior orthogonalization, the method achieves flexible nonlinear estimation and distribution correction while strictly preserving stochastic monotonicity and model identifiability. This approach establishes a semi-structured hybrid modeling paradigm that integrates high interpretability with flexibility, enabling user-defined trade-offs and quantifiable model bias. Consequently, it effectively mitigates the transparency bottleneck inherent in complex data modeling tasks.

0 citationsRead paper

Improving post-operative discharge destination prediction of geriatric patients with generative data augmentation

Apr 19, 2026

This study addresses the challenge of predicting postoperative discharge destinations in older adults, a task hindered by limited clinical data that impedes perioperative care optimization. For the first time in geriatric surgery, the authors employ Adversarial Random Forests (ARF) to generate synthetic data, which is then integrated with real-world observations using multiple imputation strategies. The combined dataset is used to train logistic regression, random forest, and TabPFN models. Results demonstrate that data augmentation substantially enhances the performance of simpler models: logistic regression accuracy improves from 0.70 to 0.81, and AUC rises from 0.85 to 0.92. In contrast, more complex models—random forest and TabPFN—maintain consistently high performance, achieving approximately 0.84 accuracy and 0.94 AUC. This work validates the efficacy and practical utility of generative machine learning approaches for clinical prediction tasks under data scarcity.

0 citationsRead paper

Explainable histomorphology-based survival prediction of glioblastoma, IDH-wildtype

Jan 16, 2026

This study addresses the challenge of automatically extracting interpretable prognostic morphological features from whole-slide images of IDH wild-type glioblastoma to predict patient survival. We propose a novel interpretable multiple instance learning framework that, for the first time, integrates a sparse autoencoder with a Cox proportional hazards model. Evaluated on 720 real-world cases, the model achieves an AUC of 0.67 in survival stratification. It identifies 24 visual patches significantly associated with survival, 21 of which were validated by neuropathologists and categorized into seven distinct histological feature classes. This approach enables an automatic mapping from raw histopathology images to clinically comprehensible morphological patterns, substantially enhancing model transparency and pathological relevance.

0 citationsRead paper

Imputation Uncertainty in Interpretable Machine Learning Methods

Dec 19, 2025

Missing values are pervasive in real-world data, yet existing interpretable machine learning (IML) methods commonly rely on single imputation, neglecting how imputation-induced uncertainty affects explanation stability—particularly confidence interval coverage probabilities. This work systematically quantifies the impact of imputation strategies on the coverage of confidence intervals for three canonical IML methods: permutation importance, partial dependence plots, and Shapley values. We compare single imputation (mean, median, regression) against multiple imputation (MICE), integrating permutation testing and nonparametric confidence interval estimation. Results show that single imputation severely underestimates variance, yielding actual 95% confidence interval coverage rates frequently below 80%. In contrast, multiple imputation restores coverage close to the nominal level across most settings, substantially enhancing the statistical reliability of IML explanations. To our knowledge, this is the first study to rigorously evaluate and quantify how imputation choices affect the inferential validity of IML outputs.

0 citationsRead paper