Institution profile

National Cancer Institute

Academic institutionnorthamerica · us
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Comparing Tobit and Two-Part Hurdle Models for Semi-Continuous Longitudinal Data with an Application to Clonal Hematopoiesis

Aug 10, 2026

This study addresses the lack of systematic guidance in choosing between Tobit and two-part hurdle models for zero-inflated semicontinuous longitudinal data. It rigorously derives, for the first time, the precise mathematical conditions under which the two models are equivalent, revealing that the Tobit model is a special case of the hurdle model when the binary component employs a probit link. Through theoretical analysis, Monte Carlo simulations, and an empirical application to clonal fraction data from the PLCO clonal hematopoiesis cohort, the authors propose a model selection criterion grounded in the plausibility of underlying assumptions. Findings indicate that while the hurdle model offers greater flexibility and robustness, the Tobit model provides more parsimonious interpretation when its assumptions hold. Both approaches yield consistent substantive conclusions in practice and substantially outperform standard linear models that ignore zero inflation, thereby corroborating established biological insights.

0 citationsRead paper

Recovering the Target Hazard Ratio Under Nonproportional Hazards Induced by an Omitted Covariate: Simulation-based Approach

Jul 29, 2026

This study addresses the bias in treatment effect estimation that arises in Cox proportional hazards models when key covariates are omitted, a problem stemming from model misspecification and potential non-proportional hazards. The authors propose a simulation-based correction method that, under the mild assumptions of a Weibull baseline hazard and unit variance for the omitted covariate, recovers an unbiased estimate of the target hazard ratio by optimizing over simulated scenarios to yield the narrowest possible confidence band for the survival curves. This approach circumvents the stringent assumptions required by existing methods and remains applicable across various censoring mechanisms and realistic parameter configurations. Simulation studies demonstrate its robustness in accurately recovering the true hazard ratio, and its practical utility and effectiveness are further validated through application to data from a phase III breast cancer clinical trial.

0 citationsRead paper

Using Importance Sampling to Estimate $p$-values in All-Subset Meta-Analysis, with Applications to Single-Cell eQTL Mapping

Apr 24, 2026

This study addresses the limitations of the ASSET method in subset-based meta-analysis, which relies on normality assumptions to compute p-values and whose analytical approximations become inaccurate under extreme tail probabilities or non-normal conditions—such as small sample sizes or low-frequency variants—while conventional Monte Carlo simulations incur prohibitive computational costs. The work presents the first systematic evaluation of ASSET’s accuracy in estimating tail p-values and introduces an efficient importance sampling (IS) algorithm that accurately estimates extremely small p-values in both independent and overlapping study designs. The proposed method maintains high precision even under non-normality and demonstrates substantial gains in computational efficiency. Its practical utility is validated through applications to the OneK1K dataset and a Korean lung cell single-cell eQTL analysis.

0 citationsRead paper

Multi-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models

Jan 06, 2026arXiv.org

Automatically mapping standardized RADS classifications from narrative radiology reports faces challenges including complex guidelines, constrained outputs, and a lack of systematic evaluation. This work introduces RXL-RADSet, the first synthetic multimodal radiology report benchmark covering ten RADS standards, comprising 1,600 radiologist-validated samples. The authors conduct a head-to-head evaluation of 41 open-source small language models (0.135–32B parameters) alongside GPT-5.2 under a unified prompting strategy. Results show that GPT-5.2 achieves 99.8% validity and 81.1% accuracy with guided prompting, while open-source models of 20–32B parameters reach approximately 99% validity and over 70% accuracy, demonstrating significant performance gains with scale. Guided prompting consistently outperforms zero-shot settings across all evaluated models.

0 citationsRead paper

The NIAID Discovery Portal: A Unified Search Engine for Infectious and Immune-Mediated Disease Datasets

Sep 16, 2025

Infectious and immune-mediated disease (IID) data are fragmented across disparate sources, lack standardized metadata schemas, and suffer from poor discoverability and reusability. Method: We developed the first unified metadata search platform specifically for the IID domain, harmonizing over 4 million dataset-level metadata records from 400+ specialized and general-purpose databases. Through format normalization, semantic integration, and construction of a domain-specific ontology, the platform enables natural-language search, predefined queries, faceted browsing, and programmatic API access. Contribution/Results: This work represents the first systematic, cross-source metadata aggregation and interoperability framework for IID data, substantially enhancing Findability, Accessibility, Interoperability, and Reusability (FAIRness). The platform is actively supporting NIH/NIAID-funded projects and global researchers in hypothesis-driven analysis, cross-cohort comparison, and secondary analysis of public datasets—thereby increasing the scientific return on investment in biomedical data infrastructure.

0 citationsRead paper
Recent publications

Latest Papers

Comparing Tobit and Two-Part Hurdle Models for Semi-Continuous Longitudinal Data with an Application to Clonal Hematopoiesis

Aug 10, 2026

This study addresses the lack of systematic guidance in choosing between Tobit and two-part hurdle models for zero-inflated semicontinuous longitudinal data. It rigorously derives, for the first time, the precise mathematical conditions under which the two models are equivalent, revealing that the Tobit model is a special case of the hurdle model when the binary component employs a probit link. Through theoretical analysis, Monte Carlo simulations, and an empirical application to clonal fraction data from the PLCO clonal hematopoiesis cohort, the authors propose a model selection criterion grounded in the plausibility of underlying assumptions. Findings indicate that while the hurdle model offers greater flexibility and robustness, the Tobit model provides more parsimonious interpretation when its assumptions hold. Both approaches yield consistent substantive conclusions in practice and substantially outperform standard linear models that ignore zero inflation, thereby corroborating established biological insights.

0 citationsRead paper

Recovering the Target Hazard Ratio Under Nonproportional Hazards Induced by an Omitted Covariate: Simulation-based Approach

Jul 29, 2026

This study addresses the bias in treatment effect estimation that arises in Cox proportional hazards models when key covariates are omitted, a problem stemming from model misspecification and potential non-proportional hazards. The authors propose a simulation-based correction method that, under the mild assumptions of a Weibull baseline hazard and unit variance for the omitted covariate, recovers an unbiased estimate of the target hazard ratio by optimizing over simulated scenarios to yield the narrowest possible confidence band for the survival curves. This approach circumvents the stringent assumptions required by existing methods and remains applicable across various censoring mechanisms and realistic parameter configurations. Simulation studies demonstrate its robustness in accurately recovering the true hazard ratio, and its practical utility and effectiveness are further validated through application to data from a phase III breast cancer clinical trial.

0 citationsRead paper

Using Importance Sampling to Estimate $p$-values in All-Subset Meta-Analysis, with Applications to Single-Cell eQTL Mapping

Apr 24, 2026

This study addresses the limitations of the ASSET method in subset-based meta-analysis, which relies on normality assumptions to compute p-values and whose analytical approximations become inaccurate under extreme tail probabilities or non-normal conditions—such as small sample sizes or low-frequency variants—while conventional Monte Carlo simulations incur prohibitive computational costs. The work presents the first systematic evaluation of ASSET’s accuracy in estimating tail p-values and introduces an efficient importance sampling (IS) algorithm that accurately estimates extremely small p-values in both independent and overlapping study designs. The proposed method maintains high precision even under non-normality and demonstrates substantial gains in computational efficiency. Its practical utility is validated through applications to the OneK1K dataset and a Korean lung cell single-cell eQTL analysis.

0 citationsRead paper

Multi-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models

Jan 06, 2026arXiv.org

Automatically mapping standardized RADS classifications from narrative radiology reports faces challenges including complex guidelines, constrained outputs, and a lack of systematic evaluation. This work introduces RXL-RADSet, the first synthetic multimodal radiology report benchmark covering ten RADS standards, comprising 1,600 radiologist-validated samples. The authors conduct a head-to-head evaluation of 41 open-source small language models (0.135–32B parameters) alongside GPT-5.2 under a unified prompting strategy. Results show that GPT-5.2 achieves 99.8% validity and 81.1% accuracy with guided prompting, while open-source models of 20–32B parameters reach approximately 99% validity and over 70% accuracy, demonstrating significant performance gains with scale. Guided prompting consistently outperforms zero-shot settings across all evaluated models.

0 citationsRead paper

The NIAID Discovery Portal: A Unified Search Engine for Infectious and Immune-Mediated Disease Datasets

Sep 16, 2025

Infectious and immune-mediated disease (IID) data are fragmented across disparate sources, lack standardized metadata schemas, and suffer from poor discoverability and reusability. Method: We developed the first unified metadata search platform specifically for the IID domain, harmonizing over 4 million dataset-level metadata records from 400+ specialized and general-purpose databases. Through format normalization, semantic integration, and construction of a domain-specific ontology, the platform enables natural-language search, predefined queries, faceted browsing, and programmatic API access. Contribution/Results: This work represents the first systematic, cross-source metadata aggregation and interoperability framework for IID data, substantially enhancing Findability, Accessibility, Interoperability, and Reusability (FAIRness). The platform is actively supporting NIH/NIAID-funded projects and global researchers in hypothesis-driven analysis, cross-cohort comparison, and secondary analysis of public datasets—thereby increasing the scientific return on investment in biomedical data infrastructure.

0 citationsRead paper