Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
研究通过临床扰动审计方法,测试医学链式思维的忠实性,发现大多数模型中链式思维与答案脱节,即使修改链式内容也不影响答案准确性。
研究通过临床扰动审计方法,测试医学链式思维的忠实性,发现大多数模型中链式思维与答案脱节,即使修改链式内容也不影响答案准确性。
本文提出ADMIL方法,通过轻量级模型PriorNet选择关键图像块,减少基础模型计算量,保持病理切片分类性能。
Existing Boolean matrix factorization methods are predominantly heuristic, sensitive to initialization, prone to local optima, and lack capabilities for model selection and uncertainty quantification, hindering their ability to uncover discrete co-variation patterns in cancer genomics. This work proposes Bayesian Boolean Matrix Factorization (BBMF), the first approach to integrate a fully conjugate Bayesian generative model into this domain. BBMF enforces strict Boolean constraints via logical AND/OR operations and leverages sparsity-inducing priors with Gibbs sampling for efficient posterior inference. The method enables automatic model selection and principled uncertainty quantification. Applied to chromosome arm-level copy number variation data from multiple myeloma patients, BBMF successfully identifies interpretable biclusters that precisely associate patient subgroups with co-varying chromosomal arms, substantially enhancing the interpretability and robustness of the inferred factors.
This work addresses the challenge of constructing valid prediction intervals for counterfactual outcomes under runtime confounding, where only a subset of confounding variables is observed in the target population. Existing conformal prediction methods often fail to achieve nominal coverage in such settings. To overcome this limitation, the paper introduces semi-parametric efficiency theory into the conformal prediction framework, integrating debiased machine learning with counterfactual modeling to produce prediction intervals that maintain valid coverage despite missing confounders. The proposed method not only resolves the coverage failure induced by unobserved confounding but also attains faster convergence rates. Empirical evaluations on multiple synthetic and semi-synthetic datasets demonstrate that the approach consistently achieves the desired coverage levels and significantly outperforms standard conformal prediction methods.
Clinical prediction from structured electronic health records (EHRs) remains challenging, with growing interest in leveraging large language models (LLMs) as surrogates despite their computational and interpretability trade-offs. Method: This work conducts the first systematic benchmark comparing traditional count-based tabular models (e.g., LightGBM, TabPFN) against emerging hybrid LLM-based pipelines—including CLMBR and table-to-text summarization—across eight clinical prediction tasks using the EHRSHOT dataset. Contribution/Results: Count-based models achieve superior balance among predictive accuracy, model interpretability, and computational efficiency, matching or exceeding hybrid LLM pipelines—especially under low-data and high-noise conditions. The study bridges a critical gap by providing the first direct empirical comparison between classical statistical methods and LLM surrogate paradigms for EHR prediction, offering evidence-based guidance and methodological insights for clinical AI deployment.
研究通过临床扰动审计方法,测试医学链式思维的忠实性,发现大多数模型中链式思维与答案脱节,即使修改链式内容也不影响答案准确性。
本文提出ADMIL方法,通过轻量级模型PriorNet选择关键图像块,减少基础模型计算量,保持病理切片分类性能。
Existing Boolean matrix factorization methods are predominantly heuristic, sensitive to initialization, prone to local optima, and lack capabilities for model selection and uncertainty quantification, hindering their ability to uncover discrete co-variation patterns in cancer genomics. This work proposes Bayesian Boolean Matrix Factorization (BBMF), the first approach to integrate a fully conjugate Bayesian generative model into this domain. BBMF enforces strict Boolean constraints via logical AND/OR operations and leverages sparsity-inducing priors with Gibbs sampling for efficient posterior inference. The method enables automatic model selection and principled uncertainty quantification. Applied to chromosome arm-level copy number variation data from multiple myeloma patients, BBMF successfully identifies interpretable biclusters that precisely associate patient subgroups with co-varying chromosomal arms, substantially enhancing the interpretability and robustness of the inferred factors.
This work addresses the challenge of constructing valid prediction intervals for counterfactual outcomes under runtime confounding, where only a subset of confounding variables is observed in the target population. Existing conformal prediction methods often fail to achieve nominal coverage in such settings. To overcome this limitation, the paper introduces semi-parametric efficiency theory into the conformal prediction framework, integrating debiased machine learning with counterfactual modeling to produce prediction intervals that maintain valid coverage despite missing confounders. The proposed method not only resolves the coverage failure induced by unobserved confounding but also attains faster convergence rates. Empirical evaluations on multiple synthetic and semi-synthetic datasets demonstrate that the approach consistently achieves the desired coverage levels and significantly outperforms standard conformal prediction methods.
Clinical prediction from structured electronic health records (EHRs) remains challenging, with growing interest in leveraging large language models (LLMs) as surrogates despite their computational and interpretability trade-offs. Method: This work conducts the first systematic benchmark comparing traditional count-based tabular models (e.g., LightGBM, TabPFN) against emerging hybrid LLM-based pipelines—including CLMBR and table-to-text summarization—across eight clinical prediction tasks using the EHRSHOT dataset. Contribution/Results: Count-based models achieve superior balance among predictive accuracy, model interpretability, and computational efficiency, matching or exceeding hybrid LLM pipelines—especially under low-data and high-noise conditions. The study bridges a critical gap by providing the first direct empirical comparison between classical statistical methods and LLM surrogate paradigms for EHR prediction, offering evidence-based guidance and methodological insights for clinical AI deployment.