Institution profile

Food and Drug Administration

Academic institutionnorthamerica · us
Official website
Research library14linked papers
Opportunities0open roles
Selected work

Representative Papers

sFRC for assessing hallucinations in medical image restoration

Mar 04, 2026

This work addresses the critical yet underexplored issue of hallucinations in deep learning–based medical image reconstruction, where outputs often appear visually realistic but contain content distortions. To tackle the lack of effective detection methods, the authors propose a scanning analysis framework based on subregion Fourier Ring Correlation (sFRC), introducing localized frequency-domain correlation analysis for the first time to medical hallucination detection. The approach supports both expert-annotated and imaging-theory–driven hallucination mapping and demonstrates broad applicability across diverse reconstruction tasks—including CT super-resolution, sparse-view CT, and undersampled MRI. In CT, it effectively identifies hallucinated regions; in MRI, its findings align closely with theoretically predicted hallucination patterns. Furthermore, the method quantifies hallucination prevalence under varying data distributions and undersampling rates, confirming its generalizability and practical utility.

2 citationsRead paper

Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation

Aug 04, 2026

Existing general-purpose image quality metrics, such as Fréchet Inception Distance (FID) and Inception Score (IS), perform poorly when evaluating synthetic histopathology images due to their reliance on ImageNet-pretrained features. This work proposes a domain-specific evaluation framework tailored for digital pathology by adapting FID and IS using a foundation model pretrained on pathological data. Additionally, it incorporates precision-recall analysis and evaluates downstream nuclei segmentation performance via AJI+ and Dice coefficients. Experimental results demonstrate that the adapted IS exhibits a strong correlation with segmentation performance (r = 0.6096, p = 0.0122), substantially outperforming the original IS (r = 0.0708). This study is the first to reveal that diversity in generated data has a greater impact on downstream task efficacy than per-image visual fidelity, thereby validating the necessity and effectiveness of domain-adapted evaluation metrics.

0 citationsRead paper

Beyond Point Estimates: Reliable Evaluation of Prediction Performance Metrics under Clustered Data

Jun 02, 2026

Current performance evaluation metrics—such as accuracy and F1 score—are typically reported as point estimates, ignoring the uncertainty induced by data clustering structures. This oversight often leads to underestimation of variability and potentially misleading model comparisons. To address this, this work proposes a unified framework that expresses a broad class of performance metrics as smooth functionals of the confusion matrix probabilities. By integrating a cluster-robust sandwich variance estimator, the framework enables valid confidence interval construction, hypothesis testing, and paired model comparison. It represents the first systematic application of cluster-robust inference to predictive performance evaluation, accommodating both binary and multiclass settings, and further provides asymptotic theory–based methods for power and sample size calculations. Simulations demonstrate that the proposed approach achieves near-nominal coverage across diverse dependence structures and substantially outperforms conventional methods that ignore clustering; real-data analyses confirm that accounting for clustering can materially alter evaluation conclusions.

0 citationsRead paper

A theory of ROC analysis of rule-out and rule-in diagnostics with applications to mammography data

Apr 25, 2026

This study addresses the challenge of effectively integrating radiologists’ assessments with AI predictions in mammographic screening to optimize rule-out and rule-in diagnostic strategies. It introduces, for the first time, a unified joint ROC theoretical framework tailored to both clinical scenarios. By modeling the dependence between physician and AI diagnostic outputs using bivariate copulas, the work theoretically derives—and empirically validates—the impact of their correlation on AUC performance: higher correlation improves rule-out efficacy in diseased populations, whereas lower correlation is preferable in non-diseased populations; conversely, for rule-in tasks, the opposite pattern holds. This framework provides a rigorous theoretical foundation and practical guidance for designing collaborative diagnostic systems that strategically leverage human–AI synergy.

0 citationsRead paper

Estimator-Aligned Prospective Sample Size Determination for Designs Using Inverse Probability of Treatment Weighting

Apr 23, 2026

This study addresses the inadequacy of conventional sample size calculations in observational studies, which neglect the uncertainty introduced by propensity score estimation and inverse probability of treatment weighting (IPTW), leading to underestimated variance and miscalibrated statistical power. The authors propose a prospective sample size determination framework aligned with IPTW estimators, integrating the propensity score model and marginal structural model into a unified estimating system via generalized estimating equations (GEE) and stacked M-estimation. This approach explicitly propagates nuisance parameter uncertainty and directly models the large-sample variance of the IPTW estimator. For the first time, it ensures consistency between sample size planning and the asymptotic variance of IPTW estimators. Leveraging pilot data for variance factor estimation and a bootstrap stabilization procedure that accounts for both internal and external variability, the method accommodates diverse outcome types. Simulations demonstrate substantially improved power calibration accuracy over traditional randomized trial–based formulas under challenging scenarios such as unstable weights, sparse outcomes, or heavy-tailed outcome distributions.

0 citationsRead paper
Recent publications

Latest Papers

Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation

Aug 04, 2026

Existing general-purpose image quality metrics, such as Fréchet Inception Distance (FID) and Inception Score (IS), perform poorly when evaluating synthetic histopathology images due to their reliance on ImageNet-pretrained features. This work proposes a domain-specific evaluation framework tailored for digital pathology by adapting FID and IS using a foundation model pretrained on pathological data. Additionally, it incorporates precision-recall analysis and evaluates downstream nuclei segmentation performance via AJI+ and Dice coefficients. Experimental results demonstrate that the adapted IS exhibits a strong correlation with segmentation performance (r = 0.6096, p = 0.0122), substantially outperforming the original IS (r = 0.0708). This study is the first to reveal that diversity in generated data has a greater impact on downstream task efficacy than per-image visual fidelity, thereby validating the necessity and effectiveness of domain-adapted evaluation metrics.

0 citationsRead paper

Beyond Point Estimates: Reliable Evaluation of Prediction Performance Metrics under Clustered Data

Jun 02, 2026

Current performance evaluation metrics—such as accuracy and F1 score—are typically reported as point estimates, ignoring the uncertainty induced by data clustering structures. This oversight often leads to underestimation of variability and potentially misleading model comparisons. To address this, this work proposes a unified framework that expresses a broad class of performance metrics as smooth functionals of the confusion matrix probabilities. By integrating a cluster-robust sandwich variance estimator, the framework enables valid confidence interval construction, hypothesis testing, and paired model comparison. It represents the first systematic application of cluster-robust inference to predictive performance evaluation, accommodating both binary and multiclass settings, and further provides asymptotic theory–based methods for power and sample size calculations. Simulations demonstrate that the proposed approach achieves near-nominal coverage across diverse dependence structures and substantially outperforms conventional methods that ignore clustering; real-data analyses confirm that accounting for clustering can materially alter evaluation conclusions.

0 citationsRead paper

A theory of ROC analysis of rule-out and rule-in diagnostics with applications to mammography data

Apr 25, 2026

This study addresses the challenge of effectively integrating radiologists’ assessments with AI predictions in mammographic screening to optimize rule-out and rule-in diagnostic strategies. It introduces, for the first time, a unified joint ROC theoretical framework tailored to both clinical scenarios. By modeling the dependence between physician and AI diagnostic outputs using bivariate copulas, the work theoretically derives—and empirically validates—the impact of their correlation on AUC performance: higher correlation improves rule-out efficacy in diseased populations, whereas lower correlation is preferable in non-diseased populations; conversely, for rule-in tasks, the opposite pattern holds. This framework provides a rigorous theoretical foundation and practical guidance for designing collaborative diagnostic systems that strategically leverage human–AI synergy.

0 citationsRead paper

Estimator-Aligned Prospective Sample Size Determination for Designs Using Inverse Probability of Treatment Weighting

Apr 23, 2026

This study addresses the inadequacy of conventional sample size calculations in observational studies, which neglect the uncertainty introduced by propensity score estimation and inverse probability of treatment weighting (IPTW), leading to underestimated variance and miscalibrated statistical power. The authors propose a prospective sample size determination framework aligned with IPTW estimators, integrating the propensity score model and marginal structural model into a unified estimating system via generalized estimating equations (GEE) and stacked M-estimation. This approach explicitly propagates nuisance parameter uncertainty and directly models the large-sample variance of the IPTW estimator. For the first time, it ensures consistency between sample size planning and the asymptotic variance of IPTW estimators. Leveraging pilot data for variance factor estimation and a bootstrap stabilization procedure that accounts for both internal and external variability, the method accommodates diverse outcome types. Simulations demonstrate substantially improved power calibration accuracy over traditional randomized trial–based formulas under challenging scenarios such as unstable weights, sparse outcomes, or heavy-tailed outcome distributions.

0 citationsRead paper

Learning, Potential, and Retention: An Approach for Evaluating Adaptive AI-Enabled Medical Devices

Apr 06, 2026

This study addresses the challenge of attributing performance changes in adaptive artificial intelligence (AI) medical devices during continuous iteration. To disentangle the effects of intrinsic model improvements from those induced by dynamic external environments, this work proposes a three-dimensional evaluation framework—comprising learning capacity, data-driven potential, and knowledge retention—and, for the first time, decouples these dimensions into independent metrics. Through case studies simulating population distribution shifts, combined with dynamic performance tracking and knowledge retention measurements, the framework demonstrates that adaptive AI systems can balance learning and stability under gradual data evolution, while exhibiting a trade-off between plasticity and stability in abrupt change scenarios. This approach enables fine-grained, regulatory-grade assessment of the evolutionary trajectory of adaptive AI systems.

0 citationsRead paper