Institution profile

Bristol Myers Squibb

Industry researchnorthamerica · us
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Assessing treatment efficacy for interval-censored endpoints using multistate semi-Markov models fit to multiple data streams

Jan 23, 2025

This study addresses the challenge of estimating treatment effects under multiple interval-censored data. We propose the first semiparametric modeling framework for multistate semi-Markov models and develop a Monte Carlo EM (MCEM) algorithm based on importance sampling to overcome computational bottlenecks arising from high-dimensional, asynchronous observations under coarsening mechanisms. Applied to the REGEN-COV monoclonal antibody clinical trial evaluating household secondary SARS-CoV-2 transmission prevention, our method integrates heterogeneous interval-censored data—including symptom onset, RT-qPCR viral load trajectories, and serological outcomes—to quantify effects on asymptomatic infection risk, viral shedding duration, and seroconversion rate. Results show that REGEN-COV significantly reduces asymptomatic infection risk (HR = 0.32), shortens median viral shedding by 4.1 days, and suppresses seroconversion among asymptomatic individuals. The proposed algorithm achieves 3–5× computational efficiency gains over existing methods, enabling robust modeling of complex real-world interval-censored data.

1 citationsRead paper

Pattern-Based Sequential Multiple Imputation for Missing Data in Clinical Trials: An Extension for Baseline-Only Early Dropout Subjects

Aug 17, 2026

This study addresses the challenge of imputing missing data for subjects with baseline-only early withdrawal in clinical trials by proposing the EPSMI-Y1 method. By integrating covariate-matched donor imputation for the first post-baseline visit with an extended pattern-mixture sequential multiple imputation framework, this approach overcomes the reliance of traditional sequential imputation on post-baseline observations and aligns effectively with treatment policy estimands. Empirical evaluations demonstrate that under informative early withdrawal mechanisms, EPSMI-Y1 significantly reduces estimation bias and improves confidence interval coverage while maintaining controlled Type I error rates. Consequently, this method provides a robust statistical solution for handling this specific missing data pattern, facilitating more reliable inference in clinical trial analyses where early dropout is non-ignorable.

0 citationsRead paper

A Deep Learning Model for Spatially Clustered Data via Differentiable Cluster Assignment

Aug 14, 2026

This study addresses nonparametric regression with unknown spatial partitions by proposing a differentiable clustering assignment mechanism that enables end-to-end joint learning of partitions and regression functions. The method integrates positional neural networks with annealed Softmax for gradient estimation and incorporates graph Laplacian regularization to prevent fragmentation, effectively handling nonlinear effects and complex error structures. Theoretically, we establish a prediction risk decomposition framework and prove that the proposed estimator achieves oracle convergence rates. Empirical evaluations demonstrate superior performance under abrupt boundaries and complex sampling designs. Collectively, this work provides a novel paradigm for modeling spatial heterogeneity in settings where partition structures are not known a priori.

0 citationsRead paper

Bridging Probabilistic LLMs and Deterministic Statistical Validation: The PROVE Multi-Agent Framework for Clinical Trial Reporting

Jul 30, 2026

This study addresses the challenge of error-prone manual verification of tables, figures, and listings (TFLs) in clinical trial reports, which often fails to detect structural or logical inconsistencies. The authors propose PROVE, a novel framework that leverages large language models (LLMs) for semantic parsing and evidence tracing of TFL content, integrated with a programmable rule engine to perform deterministic numerical and logical validation against SDTM/ADaM standards. Designed as a multi-agent architecture, PROVE combines LLM-driven semantic understanding with rule-based checks to enable auditable, configurable automated cross-verification. The approach achieves 100% accuracy under exact label matching; when confronted with linguistic variations, LLM assistance boosts recall from 0.588 to 0.993 and F1 score from 0.735 to 0.996.

0 citationsRead paper
Recent publications

Latest Papers

Pattern-Based Sequential Multiple Imputation for Missing Data in Clinical Trials: An Extension for Baseline-Only Early Dropout Subjects

Aug 17, 2026

This study addresses the challenge of imputing missing data for subjects with baseline-only early withdrawal in clinical trials by proposing the EPSMI-Y1 method. By integrating covariate-matched donor imputation for the first post-baseline visit with an extended pattern-mixture sequential multiple imputation framework, this approach overcomes the reliance of traditional sequential imputation on post-baseline observations and aligns effectively with treatment policy estimands. Empirical evaluations demonstrate that under informative early withdrawal mechanisms, EPSMI-Y1 significantly reduces estimation bias and improves confidence interval coverage while maintaining controlled Type I error rates. Consequently, this method provides a robust statistical solution for handling this specific missing data pattern, facilitating more reliable inference in clinical trial analyses where early dropout is non-ignorable.

0 citationsRead paper

A Deep Learning Model for Spatially Clustered Data via Differentiable Cluster Assignment

Aug 14, 2026

This study addresses nonparametric regression with unknown spatial partitions by proposing a differentiable clustering assignment mechanism that enables end-to-end joint learning of partitions and regression functions. The method integrates positional neural networks with annealed Softmax for gradient estimation and incorporates graph Laplacian regularization to prevent fragmentation, effectively handling nonlinear effects and complex error structures. Theoretically, we establish a prediction risk decomposition framework and prove that the proposed estimator achieves oracle convergence rates. Empirical evaluations demonstrate superior performance under abrupt boundaries and complex sampling designs. Collectively, this work provides a novel paradigm for modeling spatial heterogeneity in settings where partition structures are not known a priori.

0 citationsRead paper

Bridging Probabilistic LLMs and Deterministic Statistical Validation: The PROVE Multi-Agent Framework for Clinical Trial Reporting

Jul 30, 2026

This study addresses the challenge of error-prone manual verification of tables, figures, and listings (TFLs) in clinical trial reports, which often fails to detect structural or logical inconsistencies. The authors propose PROVE, a novel framework that leverages large language models (LLMs) for semantic parsing and evidence tracing of TFL content, integrated with a programmable rule engine to perform deterministic numerical and logical validation against SDTM/ADaM standards. Designed as a multi-agent architecture, PROVE combines LLM-driven semantic understanding with rule-based checks to enable auditable, configurable automated cross-verification. The approach achieves 100% accuracy under exact label matching; when confronted with linguistic variations, LLM assistance boosts recall from 0.588 to 0.993 and F1 score from 0.735 to 0.996.

0 citationsRead paper

Robust estimation of causal dose-response relationship using exposure data with dose as an instrumental variable

Aug 06, 2025

This study addresses bias in dose–response estimation in randomized trials arising from measurement error in drug exposure or unobserved confounding. We propose a robust causal inference method leveraging dose itself as a natural instrumental variable (IV), integrating control-function and ANCOVA-based adjustment. Unlike conventional approaches, our method does not require correct specification of the exposure–outcome model and remains consistent and asymptotically normal even under model misspecification. Theoretical analysis establishes its robustness to unobserved confounding, while simulations demonstrate excellent finite-sample performance. Applied to a CAR-T cell therapy clinical trial, it significantly improves accuracy in identifying the optimal dose. Our key contribution is the first formalization of dose as a built-in IV, enabling a model-agnostic, misspecification-robust framework for causal dose–response estimation—balancing interpretability with statistical rigor.

0 citationsRead paper