Institution profile

Merck & Co., Inc.

Industry researchnorthamerica · us
Official website
Research library34linked papers
Opportunities0open roles
Selected work

Representative Papers

Bayesian Joint Additive Factor Models for Multiview Learning

Jun 02, 2024arXiv.org

To address challenges in multi-omics and other multi-view data—including difficulty modeling cross-view dependencies, strong signal heterogeneity, and insufficient interpretability and uncertainty quantification—this paper proposes JAFAR, a joint Bayesian factor model. Methodologically, JAFAR introduces the Dependency-Cumulative Shrinkage Prior (D-CUSP), which jointly characterizes shared and view-specific latent factor structures while ensuring parameter identifiability. It integrates Bayesian nonparametrics, structured additive designs, partially collapsed Gibbs sampling, and flexible distributional extensions—accommodating non-Gaussian features and survival outcomes. In an application to preterm birth prediction, JAFAR jointly analyzes immunomic, metabolomic, and proteomic data, achieving statistically significant improvements over state-of-the-art methods. The model enables interpretable feature selection and principled uncertainty quantification. An open-source R package implementing JAFAR is publicly available.

1 citationsRead paper

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

Aug 06, 2026

This work addresses the challenge of reaction yield prediction, which is hindered by scarce labeled data, the vast and sparse reaction space, and the inability of existing representations to capture complex chemical transformations. To overcome these limitations, the authors propose RxnCLF, a self-supervised contrastive learning framework that introduces a novel condensed reaction graph (CRG) integrating both reactant and product information. By leveraging graph neural networks, RxnCLF learns explicit and interpretable transformation structures and models chemical reactions within a unified continuous latent space. The method significantly outperforms current graph- and sequence-based models on multiple yield prediction benchmarks, demonstrating substantial improvements in R² scores, and exhibits strong generalization capabilities on downstream tasks such as regioselectivity and enantioselectivity prediction, offering a new paradigm toward a general-purpose foundation model for chemical reactions.

0 citationsRead paper

Two-Stage Design with Sample Size Re-estimation Using gsDesign

Aug 04, 2026

This study addresses the challenges in early-phase clinical trials arising from uncertainty in effect size and nuisance parameters, which often lead to inaccurate sample size planning and over-enrollment in group sequential designs. The authors employ a two-stage group sequential framework that dynamically adjusts the sample size at interim analysis based on conditional power, and systematically compare—using the gsDesign R package—the expected sample size and statistical power of conditional power–based designs against more conservative group sequential approaches. Findings indicate that conditional power designs offer only marginal benefits under specific scenarios, whereas conservative designs generally demonstrate superior efficiency, robustness, and the added advantage of avoiding premature disclosure of interim treatment efficacy. These results provide empirical evidence and practical guidance for selecting appropriate group sequential designs in clinical trial settings.

0 citationsRead paper

Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics

Jul 31, 2026

This study addresses the challenge of information silos in fault detection and diagnosis (FDD) for variable air volume (VAV) HVAC systems, which arise from data heterogeneity and semantic ambiguity. To this end, the authors propose FDD-ON, the first structured semantic framework specifically designed for this domain. Built upon modular ontology engineering and knowledge graph techniques, FDD-ON formally models system components, fault types, symptoms, impacts, and their causal relationships, providing a unified controlled vocabulary and a comprehensive knowledge base. The framework enables machine-interpretable diagnostic reasoning and enhances system interoperability. Experimental validation on public datasets demonstrates that FDD-ON serves as a foundational semantic infrastructure for developing transparent, scalable, and AI-driven FDD applications.

0 citationsRead paper

Conformal Prediction for Regression with Clipped Outcomes

Jul 22, 2026

This work addresses the challenge of conformal prediction in regression settings where the response variable is subject to two-sided truncation. Existing methods struggle to simultaneously achieve marginal and conditional coverage, often failing to provide valid conditional coverage—particularly on easily predictable instances. To overcome this limitation, the authors introduce a novel nonconformity score tailored to truncated data and propose two calibration strategies: one ensuring tight marginal coverage, and another employing a two-stage mechanism that prioritizes conditional coverage, thereby exposing the inherent limitations of marginal coverage in truncation scenarios. Within the conformal prediction framework, the proposed score leverages the structure of truncated observations to deliver finite-sample theoretical coverage guarantees and, under model consistency, attains oracle-like asymptotic performance, substantially outperforming naive adaptations of existing methods.

0 citationsRead paper
Recent publications

Latest Papers

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

Aug 06, 2026

This work addresses the challenge of reaction yield prediction, which is hindered by scarce labeled data, the vast and sparse reaction space, and the inability of existing representations to capture complex chemical transformations. To overcome these limitations, the authors propose RxnCLF, a self-supervised contrastive learning framework that introduces a novel condensed reaction graph (CRG) integrating both reactant and product information. By leveraging graph neural networks, RxnCLF learns explicit and interpretable transformation structures and models chemical reactions within a unified continuous latent space. The method significantly outperforms current graph- and sequence-based models on multiple yield prediction benchmarks, demonstrating substantial improvements in R² scores, and exhibits strong generalization capabilities on downstream tasks such as regioselectivity and enantioselectivity prediction, offering a new paradigm toward a general-purpose foundation model for chemical reactions.

0 citationsRead paper

Two-Stage Design with Sample Size Re-estimation Using gsDesign

Aug 04, 2026

This study addresses the challenges in early-phase clinical trials arising from uncertainty in effect size and nuisance parameters, which often lead to inaccurate sample size planning and over-enrollment in group sequential designs. The authors employ a two-stage group sequential framework that dynamically adjusts the sample size at interim analysis based on conditional power, and systematically compare—using the gsDesign R package—the expected sample size and statistical power of conditional power–based designs against more conservative group sequential approaches. Findings indicate that conditional power designs offer only marginal benefits under specific scenarios, whereas conservative designs generally demonstrate superior efficiency, robustness, and the added advantage of avoiding premature disclosure of interim treatment efficacy. These results provide empirical evidence and practical guidance for selecting appropriate group sequential designs in clinical trial settings.

0 citationsRead paper

Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics

Jul 31, 2026

This study addresses the challenge of information silos in fault detection and diagnosis (FDD) for variable air volume (VAV) HVAC systems, which arise from data heterogeneity and semantic ambiguity. To this end, the authors propose FDD-ON, the first structured semantic framework specifically designed for this domain. Built upon modular ontology engineering and knowledge graph techniques, FDD-ON formally models system components, fault types, symptoms, impacts, and their causal relationships, providing a unified controlled vocabulary and a comprehensive knowledge base. The framework enables machine-interpretable diagnostic reasoning and enhances system interoperability. Experimental validation on public datasets demonstrates that FDD-ON serves as a foundational semantic infrastructure for developing transparent, scalable, and AI-driven FDD applications.

0 citationsRead paper

Conformal Prediction for Regression with Clipped Outcomes

Jul 22, 2026

This work addresses the challenge of conformal prediction in regression settings where the response variable is subject to two-sided truncation. Existing methods struggle to simultaneously achieve marginal and conditional coverage, often failing to provide valid conditional coverage—particularly on easily predictable instances. To overcome this limitation, the authors introduce a novel nonconformity score tailored to truncated data and propose two calibration strategies: one ensuring tight marginal coverage, and another employing a two-stage mechanism that prioritizes conditional coverage, thereby exposing the inherent limitations of marginal coverage in truncation scenarios. Within the conformal prediction framework, the proposed score leverages the structure of truncated observations to deliver finite-sample theoretical coverage guarantees and, under model consistency, attains oracle-like asymptotic performance, substantially outperforming naive adaptations of existing methods.

0 citationsRead paper

Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks

Jul 13, 2026

Current whole-slide image (WSI) multimodal visual question answering (VQA) benchmarks are severely compromised by patient- and institution-level data leakage, leading to inflated estimates of model reasoning capabilities. This work presents the first systematic audit of dual leakage issues in publicly available WSI VQA datasets. Through identifier tracing, linear separability analysis in feature space, and performance comparisons between leaked and clean samples, we reveal case overlap rates of 92.3%–100% in TCGA-derived benchmarks and demonstrate that leakage signals can be linearly decoded from features extracted by foundation models. Our findings indicate that reported high accuracies primarily stem from memorization of leakage artifacts rather than genuine multimodal reasoning. Based on these insights, we propose concrete guidelines for constructing contamination-free evaluation protocols.

0 citationsRead paper