Institution profile

Dana-Farber Cancer Institute

Academic institutionnorthamerica · us
Official website
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

A Bayesian Boolean Matrix Factorization with Application to Copy Number Analysis in Cancer

Jun 16, 2026

Existing Boolean matrix factorization methods are predominantly heuristic, sensitive to initialization, prone to local optima, and lack capabilities for model selection and uncertainty quantification, hindering their ability to uncover discrete co-variation patterns in cancer genomics. This work proposes Bayesian Boolean Matrix Factorization (BBMF), the first approach to integrate a fully conjugate Bayesian generative model into this domain. BBMF enforces strict Boolean constraints via logical AND/OR operations and leverages sparsity-inducing priors with Gibbs sampling for efficient posterior inference. The method enables automatic model selection and principled uncertainty quantification. Applied to chromosome arm-level copy number variation data from multiple myeloma patients, BBMF successfully identifies interpretable biclusters that precisely associate patient subgroups with co-varying chromosomal arms, substantially enhancing the interpretability and robustness of the inferred factors.

0 citationsRead paper

Debiased Machine Learning for Conformal Prediction of Counterfactual Outcomes Under Runtime Confounding

Apr 04, 2026

This work addresses the challenge of constructing valid prediction intervals for counterfactual outcomes under runtime confounding, where only a subset of confounding variables is observed in the target population. Existing conformal prediction methods often fail to achieve nominal coverage in such settings. To overcome this limitation, the paper introduces semi-parametric efficiency theory into the conformal prediction framework, integrating debiased machine learning with counterfactual modeling to produce prediction intervals that maintain valid coverage despite missing confounders. The proposed method not only resolves the coverage failure induced by unobserved confounding but also attains faster convergence rates. Empirical evaluations on multiple synthetic and semi-synthetic datasets demonstrate that the approach consistently achieves the desired coverage levels and significantly outperforms standard conformal prediction methods.

0 citationsRead paper

Count-Based Approaches Remain Strong: A Benchmark Against Transformer and LLM Pipelines on Structured EHR

Nov 01, 2025

Clinical prediction from structured electronic health records (EHRs) remains challenging, with growing interest in leveraging large language models (LLMs) as surrogates despite their computational and interpretability trade-offs. Method: This work conducts the first systematic benchmark comparing traditional count-based tabular models (e.g., LightGBM, TabPFN) against emerging hybrid LLM-based pipelines—including CLMBR and table-to-text summarization—across eight clinical prediction tasks using the EHRSHOT dataset. Contribution/Results: Count-based models achieve superior balance among predictive accuracy, model interpretability, and computational efficiency, matching or exceeding hybrid LLM pipelines—especially under low-data and high-noise conditions. The study bridges a critical gap by providing the first direct empirical comparison between classical statistical methods and LLM surrogate paradigms for EHR prediction, offering evidence-based guidance and methodological insights for clinical AI deployment.

0 citationsRead paper
Recent publications

Latest Papers

A Bayesian Boolean Matrix Factorization with Application to Copy Number Analysis in Cancer

Jun 16, 2026

Existing Boolean matrix factorization methods are predominantly heuristic, sensitive to initialization, prone to local optima, and lack capabilities for model selection and uncertainty quantification, hindering their ability to uncover discrete co-variation patterns in cancer genomics. This work proposes Bayesian Boolean Matrix Factorization (BBMF), the first approach to integrate a fully conjugate Bayesian generative model into this domain. BBMF enforces strict Boolean constraints via logical AND/OR operations and leverages sparsity-inducing priors with Gibbs sampling for efficient posterior inference. The method enables automatic model selection and principled uncertainty quantification. Applied to chromosome arm-level copy number variation data from multiple myeloma patients, BBMF successfully identifies interpretable biclusters that precisely associate patient subgroups with co-varying chromosomal arms, substantially enhancing the interpretability and robustness of the inferred factors.

0 citationsRead paper

Debiased Machine Learning for Conformal Prediction of Counterfactual Outcomes Under Runtime Confounding

Apr 04, 2026

This work addresses the challenge of constructing valid prediction intervals for counterfactual outcomes under runtime confounding, where only a subset of confounding variables is observed in the target population. Existing conformal prediction methods often fail to achieve nominal coverage in such settings. To overcome this limitation, the paper introduces semi-parametric efficiency theory into the conformal prediction framework, integrating debiased machine learning with counterfactual modeling to produce prediction intervals that maintain valid coverage despite missing confounders. The proposed method not only resolves the coverage failure induced by unobserved confounding but also attains faster convergence rates. Empirical evaluations on multiple synthetic and semi-synthetic datasets demonstrate that the approach consistently achieves the desired coverage levels and significantly outperforms standard conformal prediction methods.

0 citationsRead paper

Count-Based Approaches Remain Strong: A Benchmark Against Transformer and LLM Pipelines on Structured EHR

Nov 01, 2025

Clinical prediction from structured electronic health records (EHRs) remains challenging, with growing interest in leveraging large language models (LLMs) as surrogates despite their computational and interpretability trade-offs. Method: This work conducts the first systematic benchmark comparing traditional count-based tabular models (e.g., LightGBM, TabPFN) against emerging hybrid LLM-based pipelines—including CLMBR and table-to-text summarization—across eight clinical prediction tasks using the EHRSHOT dataset. Contribution/Results: Count-based models achieve superior balance among predictive accuracy, model interpretability, and computational efficiency, matching or exceeding hybrid LLM pipelines—especially under low-data and high-noise conditions. The study bridges a critical gap by providing the first direct empirical comparison between classical statistical methods and LLM surrogate paradigms for EHR prediction, offering evidence-based guidance and methodological insights for clinical AI deployment.

0 citationsRead paper