Institution profile

Hospital for Sick Children

Academic institutionnorthamerica · ca
Official website
Research library34linked papers
Opportunities0open roles
Selected work

Representative Papers

Scalable and Cost-Efficient de Novo Template-Based Molecular Generation

Jun 10, 2025arXiv.org

This work addresses three key challenges in template-guided molecular generation: high synthetic cost, difficulty in scaling the building block library, and underutilization of small fragments. We propose a recursive, cost-guided generative framework based on Generative Flow Networks (GFlowNets). Methodologically, we design a backward policy network coupled with an auxiliary synthetic cost predictor, introduce a dynamic building block library that reuses intermediate molecular states, and employ a penalty mechanism to balance exploration and exploitation. Our key contribution lies in explicitly embedding synthetic cost into the generative process, enabling end-to-end differentiable optimization; the dynamic library mechanism markedly improves both diversity and efficiency—especially with limited building block sets. On standard templated molecular generation benchmarks, our approach generates higher-quality, more diverse molecules at lower synthetic cost, achieving state-of-the-art performance.

2 citationsRead paper

Neural Collision Detection for Multi-arm Laparoscopy Surgical Robots Through Learning-from-Simulation

Jan 21, 2026

This study addresses the challenge of collision risk and minimum distance estimation in multi-arm laparoscopic surgical robots by proposing a real-time collision detection framework that integrates analytical modeling, 3D physics simulation, and deep learning. The approach constructs an analytical geometric model of a 7-degree-of-freedom robotic arm and generates diverse configuration data to train a deep neural network that predicts inter-arm minimum distances using joint states and relative poses as inputs. By uniquely combining analytical modeling with simulation-driven deep learning, the method achieves high-accuracy spatial relationship generalization and enables real-time collision warnings. Experimental results demonstrate the model’s accuracy and generalization capability, yielding a mean absolute error of 282.2 mm and an R² score of 0.85 in minimum distance prediction.

1 citationsRead paper

Early Pregnancy Treatment Decisions: Designing Perinatal Pharmacoepidemiology Studies using Real-World Data

Aug 11, 2026

This study addresses methodological challenges in perinatal pharmacologic research—such as left and right censoring, competing events, and gestational age heterogeneity—that often introduce bias in causal effect estimation, particularly when evaluating decisions to modify preconception treatment regimens during early pregnancy. Leveraging real-world data, this work extends the target trial emulation framework to the context of pre-pregnancy medication changes and innovatively proposes a gestational age–anchored definition of time zero tailored to early-pregnancy therapeutic decisions. The approach systematically corrects for selection bias and immortal time bias. By establishing a generalizable methodological paradigm, this research enhances the scientific rigor and feasibility of pharmacoepidemiologic studies assessing the safety and effectiveness of medications—such as those for type 2 diabetes—during the periconceptional and prenatal periods.

0 citationsRead paper

LiNC: Lightweight Noise Correction via Per-Sample Trust and Gaussian Mixture Modeling

Aug 04, 2026

This work addresses label noise in medical image datasets arising from annotator disagreement, errors, and ambiguous samples by proposing LiNC, a lightweight noise correction method. LiNC introduces a learnable trust parameter for each sample during standard training, dynamically blending the observed label with the model’s prediction via a convex combination. A three-component Gaussian mixture model clusters these trust values to distinguish clean, ambiguous, and noisy samples, enabling a two-stage soft-then-hard correction strategy. Without requiring auxiliary networks or complex pipelines, LiNC achieves significant improvements in classification accuracy and highly precise mislabel detection across ten 2D datasets in MedMNISTv2 under noise rates as high as 50%, while incurring negligible additional training overhead and only linearly increasing memory usage.

0 citationsRead paper
Recent publications

Latest Papers

Early Pregnancy Treatment Decisions: Designing Perinatal Pharmacoepidemiology Studies using Real-World Data

Aug 11, 2026

This study addresses methodological challenges in perinatal pharmacologic research—such as left and right censoring, competing events, and gestational age heterogeneity—that often introduce bias in causal effect estimation, particularly when evaluating decisions to modify preconception treatment regimens during early pregnancy. Leveraging real-world data, this work extends the target trial emulation framework to the context of pre-pregnancy medication changes and innovatively proposes a gestational age–anchored definition of time zero tailored to early-pregnancy therapeutic decisions. The approach systematically corrects for selection bias and immortal time bias. By establishing a generalizable methodological paradigm, this research enhances the scientific rigor and feasibility of pharmacoepidemiologic studies assessing the safety and effectiveness of medications—such as those for type 2 diabetes—during the periconceptional and prenatal periods.

0 citationsRead paper

LiNC: Lightweight Noise Correction via Per-Sample Trust and Gaussian Mixture Modeling

Aug 04, 2026

This work addresses label noise in medical image datasets arising from annotator disagreement, errors, and ambiguous samples by proposing LiNC, a lightweight noise correction method. LiNC introduces a learnable trust parameter for each sample during standard training, dynamically blending the observed label with the model’s prediction via a convex combination. A three-component Gaussian mixture model clusters these trust values to distinguish clean, ambiguous, and noisy samples, enabling a two-stage soft-then-hard correction strategy. Without requiring auxiliary networks or complex pipelines, LiNC achieves significant improvements in classification accuracy and highly precise mislabel detection across ten 2D datasets in MedMNISTv2 under noise rates as high as 50%, while incurring negligible additional training overhead and only linearly increasing memory usage.

0 citationsRead paper

Bayesian hierarchical bootstrap framework for causal subgroup estimation with a time-to-event outcome

Aug 01, 2026

This study addresses the challenges of estimating causal treatment effects within predefined subgroups, where sparse samples, unstable covariate distributions, and right censoring often lead to high bias and uncertainty. The authors propose a novel approach that extends hierarchical Bayesian bootstrap (HBB) to time-to-event causal subgroup analysis with right-censored data. By integrating a Bayesian accelerated failure time model with a nonparametric hierarchical prior, the method models subgroup-specific baseline covariate distributions while borrowing strength across subgroups. Joint uncertainty from both the survival model and covariate distribution is propagated via the posterior g-formula. Simulation studies demonstrate that the proposed method substantially improves estimation stability and accuracy across varying levels of sparsity and censoring, outperforming existing alternatives such as Bayesian additive regression trees.

0 citationsRead paper

Calculating the Expected Value of Sample Information accounting for missing data

Jul 28, 2026

This study addresses the challenge that existing expected value of sample information (EVSI) computation methods struggle to accommodate missing data, a common issue in real-world research. The authors propose a novel EVSI estimation approach tailored for complex missingness mechanisms—including missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR)—by integrating individual-level data simulation, multiple imputation, and nonparametric regression. Notably, this work extends the EVSI framework to settings with non-ignorable missingness for the first time. Building on the EVSI loss, the study further introduces a new sample size determination strategy. Findings demonstrate that missing data substantially reduce EVSI, and achieving the EVSI attainable under complete data requires markedly larger sample sizes than conventional calculations suggest, thereby enhancing the practical applicability of health economic decision models.

0 citationsRead paper