Institution profile

Novo Nordisk

Industry researcheurope · dk
Official website
Research library23linked papers
Opportunities0open roles
Selected work

Representative Papers

Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory

Jul 30, 2026

This work addresses the challenge of efficiently selecting high-potential candidate molecules under a limited oracle query budget in molecular optimization. The authors propose a plug-in short-term graph memory mechanism that, without altering the generator architecture, incrementally constructs a graph neural network surrogate model from previously evaluated molecules to prescreen candidates and prioritize those with higher predicted utility for oracle evaluation. This approach incurs no additional oracle calls yet significantly enhances optimization efficiency. Experimental results demonstrate that, under a stringent budget of only 1,000 oracle queries, the method consistently outperforms the original strategy across four fragment-based generators, substantially improving the average top-10 score without any observed performance degradation, while also revealing systematic interactions between the generator’s exploration–exploitation behavior and surrogate-guided selection.

0 citationsRead paper

Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble Modeling

Jul 23, 2026

Traditional molecular property prediction methods typically rely on a single conformation, overlooking the fact that molecules exist as ensembles of conformers in solution. This work proposes EnsembleEGNN, a novel model that represents the entire conformational ensemble as a unified embedding by encoding individual conformers with a shared equivariant graph neural network (EGNN) and aggregating ensemble information via a set attention mechanism. The model is pretrained through a multi-task self-supervised strategy involving masked token recovery, noisy coordinate reconstruction, and pairwise distance reconstruction, and is jointly optimized with a BERT sequence encoder. Evaluated on the CREMP-CycPeptMPDB dataset, EnsembleEGNN achieves an R² of 0.538 (Pearson r = 0.737), substantially outperforming existing baselines and demonstrating the efficacy and superiority of explicit conformational ensemble modeling for cyclic peptide property prediction.

0 citationsRead paper

Optimizing Large Language Models for Causality Assessment in Pharmacovigilance: Developing a Performance Metric as Objective for Bayesian Hyperparameter Optimization

Jul 04, 2026

This study addresses the suboptimal performance of large language models (LLMs) in causality assessment for automated pharmacovigilance and the absence of effective methods for optimizing inference hyperparameters such as temperature. To this end, the authors propose a Gaussian process–based Bayesian optimization framework that systematically tunes temperature for LLM-based causal inference, incorporating a novel weighted consistency metric—particularly the Entropy-Weighted Agreement Consistency Score (EWACS). Evaluated on individual case safety reports from FAERS using GPT-5.2, chain-of-thought prompting, and four consistency measures, the approach significantly improves agreement between model predictions and expert judgments from 45.0% to 72.0%, with a 42.9-percentage-point gain in the “suspected” category. These results demonstrate that optimal temperature is highly context-dependent, precluding a universal setting, and substantially enhance the practical utility of LLMs in regulatory pharmacovigilance applications.

0 citationsRead paper

One-step Outcome Imputation: An Alternative to Multiple Imputation

Jun 05, 2026

In randomized controlled trials, methods such as reference-based multiple imputation often yield invalid standard error estimates due to violations of Rubin’s rules. This work proposes a one-step estimation framework that directly leverages the influence function of the target treatment effect’s imputation model to construct an asymptotically efficient estimator, thereby circumventing the need for multiple imputation. The approach accommodates various settings—including reference-based imputation and imputation under intercurrent event dependence—while ensuring consistency of the treatment effect estimate and achieving asymptotic validity of standard errors. By doing so, it substantially enhances inferential reliability and reduces computational burden compared to conventional multiple imputation strategies.

0 citationsRead paper

Restricted mean time lost for survival and competing risks data using mets in R

May 28, 2026

This work addresses the lack of efficient and unified tools for estimating restricted mean survival time (RMST) and restricted mean time lost (RMTL) under both standard survival and competing risks settings, particularly the absence of methods for cause-specific RMTL and its causal effects. Building upon the mets R package, we present the first implementation capable of simultaneously computing RMST/RMTL and their standard errors over the full time range. Our approach integrates inverse probability censoring weighting (IPCW), G-computation, and influence function theory to enable nonparametric estimation, regression modeling, standardization, and inference for average treatment effects (ATE). The proposed method achieves linear time complexity, substantially enhancing computational efficiency and statistical flexibility for large-scale data analysis.

0 citationsRead paper
Recent publications

Latest Papers

Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory

Jul 30, 2026

This work addresses the challenge of efficiently selecting high-potential candidate molecules under a limited oracle query budget in molecular optimization. The authors propose a plug-in short-term graph memory mechanism that, without altering the generator architecture, incrementally constructs a graph neural network surrogate model from previously evaluated molecules to prescreen candidates and prioritize those with higher predicted utility for oracle evaluation. This approach incurs no additional oracle calls yet significantly enhances optimization efficiency. Experimental results demonstrate that, under a stringent budget of only 1,000 oracle queries, the method consistently outperforms the original strategy across four fragment-based generators, substantially improving the average top-10 score without any observed performance degradation, while also revealing systematic interactions between the generator’s exploration–exploitation behavior and surrogate-guided selection.

0 citationsRead paper

Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble Modeling

Jul 23, 2026

Traditional molecular property prediction methods typically rely on a single conformation, overlooking the fact that molecules exist as ensembles of conformers in solution. This work proposes EnsembleEGNN, a novel model that represents the entire conformational ensemble as a unified embedding by encoding individual conformers with a shared equivariant graph neural network (EGNN) and aggregating ensemble information via a set attention mechanism. The model is pretrained through a multi-task self-supervised strategy involving masked token recovery, noisy coordinate reconstruction, and pairwise distance reconstruction, and is jointly optimized with a BERT sequence encoder. Evaluated on the CREMP-CycPeptMPDB dataset, EnsembleEGNN achieves an R² of 0.538 (Pearson r = 0.737), substantially outperforming existing baselines and demonstrating the efficacy and superiority of explicit conformational ensemble modeling for cyclic peptide property prediction.

0 citationsRead paper

Optimizing Large Language Models for Causality Assessment in Pharmacovigilance: Developing a Performance Metric as Objective for Bayesian Hyperparameter Optimization

Jul 04, 2026

This study addresses the suboptimal performance of large language models (LLMs) in causality assessment for automated pharmacovigilance and the absence of effective methods for optimizing inference hyperparameters such as temperature. To this end, the authors propose a Gaussian process–based Bayesian optimization framework that systematically tunes temperature for LLM-based causal inference, incorporating a novel weighted consistency metric—particularly the Entropy-Weighted Agreement Consistency Score (EWACS). Evaluated on individual case safety reports from FAERS using GPT-5.2, chain-of-thought prompting, and four consistency measures, the approach significantly improves agreement between model predictions and expert judgments from 45.0% to 72.0%, with a 42.9-percentage-point gain in the “suspected” category. These results demonstrate that optimal temperature is highly context-dependent, precluding a universal setting, and substantially enhance the practical utility of LLMs in regulatory pharmacovigilance applications.

0 citationsRead paper

One-step Outcome Imputation: An Alternative to Multiple Imputation

Jun 05, 2026

In randomized controlled trials, methods such as reference-based multiple imputation often yield invalid standard error estimates due to violations of Rubin’s rules. This work proposes a one-step estimation framework that directly leverages the influence function of the target treatment effect’s imputation model to construct an asymptotically efficient estimator, thereby circumventing the need for multiple imputation. The approach accommodates various settings—including reference-based imputation and imputation under intercurrent event dependence—while ensuring consistency of the treatment effect estimate and achieving asymptotic validity of standard errors. By doing so, it substantially enhances inferential reliability and reduces computational burden compared to conventional multiple imputation strategies.

0 citationsRead paper

Restricted mean time lost for survival and competing risks data using mets in R

May 28, 2026

This work addresses the lack of efficient and unified tools for estimating restricted mean survival time (RMST) and restricted mean time lost (RMTL) under both standard survival and competing risks settings, particularly the absence of methods for cause-specific RMTL and its causal effects. Building upon the mets R package, we present the first implementation capable of simultaneously computing RMST/RMTL and their standard errors over the full time range. Our approach integrates inverse probability censoring weighting (IPCW), G-computation, and influence function theory to enable nonparametric estimation, regression modeling, standardization, and inference for average treatment effects (ATE). The proposed method achieves linear time complexity, substantially enhancing computational efficiency and statistical flexibility for large-scale data analysis.

0 citationsRead paper