Institution profile

Roche

Industry researcheurope · ch
Official website
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)

Jan 08, 2026arXiv.org

This work addresses the inefficiency and inconsistency of manual annotation in traditional pharmaceutical video processing, which struggles to leverage multimodal information—particularly in large-scale, long-form videos such as clinical trial interviews. The authors propose an end-to-end framework for automatically generating highlight clips by integrating vision-language models (VLMs) and audio-language models (ALMs), enhanced with role-based prompting to enable personalized editing tailored to marketing, training, and regulatory scenarios. Key innovations include a reproducible Cut & Merge algorithm ensuring audiovisual synchronization and smooth transitions, a role-prompt-driven personalization mechanism, and a highly efficient, low-cost pipeline. Evaluated on the Video MME benchmark and a dataset of 16,159 pharmaceutical videos, the system achieves a 3–4× speedup and 4× cost reduction compared to baselines, while outperforming advanced models like Gemini 2.5 Pro in both coherence (0.348) and informativeness (0.721).

1 citationsRead paper

Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform

Jan 08, 2026arXiv.org

This work addresses the challenge of efficiently processing long videos in pharmaceutical industrial settings, where GPU resources, latency, and cost are tightly constrained. The authors propose the first industrial-scale multimodal generative AI framework tailored for the pharmaceutical domain, integrating scaled dot-product attention (SDPA), a novel multimodal fusion strategy, and keyframe extraction to achieve 3–8× inference acceleration on commodity GPUs. Systematic evaluation on a large-scale dataset—comprising over 200,000 PDFs, 25,000 long videos, and 888 multilingual audio samples—demonstrates that the multimodal approach significantly outperforms unimodal baselines on 8 out of 12 tasks, with particularly strong gains in video-length-dependent scenarios. The study further identifies four critical bottlenecks: multimodal fusion design, temporal reasoning limits, attention mechanism trade-offs, and video segmentation strategies.

1 citationsRead paper

Diffeomorphic Optimization

Jul 01, 2026

This work addresses differentiable optimization on low-dimensional data manifolds embedded in high-dimensional spaces, where conventional gradient descent often deviates from the manifold and struggles with non-convex loss landscapes. The authors propose a novel approach that leverages diffusion and flow models to construct a diffeomorphic mapping, pulling the manifold back to a simple base space for optimization. Using tools from differential geometry, they prove this procedure is equivalent to Riemannian gradient descent, inherently preserving trajectories on the manifold. This is the first integration of diffeomorphic mappings with Riemannian optimization, extended to the Lie groups SO(3) and SE(3), yielding an automatic differentiation–compatible SO(3) gradient and a generalized adjoint-state backpropagation for Lie group ODE solvers. In protein design tasks, FrameFlow achieves a 91.3% secondary structure targeting success rate (versus 63.3% baseline), doubles the peptide binding affinity optimization speed compared to OC-Flow, and significantly reduces Rosetta energy by thousands of units.

0 citationsRead paper

OphthaDT: Generative Digital Twins for Forecasting Visual Acuity Trajectories in Ophthalmology

Jun 20, 2026

This study addresses the challenge of long-term visual trajectory prediction in ophthalmic precision medicine, hindered by fragmented multimodal clinical data. It proposes the first large language model (LLM)-driven digital twin system for ophthalmology, which transforms longitudinal medical histories from 3,220 patients into structured clinical narratives. The system effectively handles irregularly sampled time-series data without requiring imputation, demonstrating enhanced modeling capacity under high clinical variability. In neovascular age-related macular degeneration (nAMD) prediction, it achieves a 6.0% reduction in mean absolute error (MAE) compared to existing baselines. For diabetic macular edema (DME), it outperforms Random Forest and XGBoost by 2.6% and 6.9% in MAE, respectively. These results validate the innovation and superiority of LLM-powered generative digital twins in longitudinal ophthalmic forecasting.

0 citationsRead paper

Communicating results in trials with multiple hypotheses or adaptive design features

May 05, 2026

This study addresses the underappreciated challenge of estimation and communication following multiplicity adjustment within the frequentist framework in complex clinical trials, where multiple endpoints, interim data looks, or group comparisons often introduce estimation bias and complicate interpretation, thereby undermining transparency in benefit–risk assessment. By integrating advanced methodologies such as adaptive designs and graphical approaches to multiple testing, the work illustrates through concrete examples the limitations of current strategies in conveying trial results meaningfully. The research underscores the need to critically reevaluate prevailing practices and foster interdisciplinary dialogue to enhance both the accuracy of effect estimation and the clarity of result communication, ultimately informing future methodological standards and regulatory guidance.

0 citationsRead paper
Recent publications

Latest Papers

Diffeomorphic Optimization

Jul 01, 2026

This work addresses differentiable optimization on low-dimensional data manifolds embedded in high-dimensional spaces, where conventional gradient descent often deviates from the manifold and struggles with non-convex loss landscapes. The authors propose a novel approach that leverages diffusion and flow models to construct a diffeomorphic mapping, pulling the manifold back to a simple base space for optimization. Using tools from differential geometry, they prove this procedure is equivalent to Riemannian gradient descent, inherently preserving trajectories on the manifold. This is the first integration of diffeomorphic mappings with Riemannian optimization, extended to the Lie groups SO(3) and SE(3), yielding an automatic differentiation–compatible SO(3) gradient and a generalized adjoint-state backpropagation for Lie group ODE solvers. In protein design tasks, FrameFlow achieves a 91.3% secondary structure targeting success rate (versus 63.3% baseline), doubles the peptide binding affinity optimization speed compared to OC-Flow, and significantly reduces Rosetta energy by thousands of units.

0 citationsRead paper

OphthaDT: Generative Digital Twins for Forecasting Visual Acuity Trajectories in Ophthalmology

Jun 20, 2026

This study addresses the challenge of long-term visual trajectory prediction in ophthalmic precision medicine, hindered by fragmented multimodal clinical data. It proposes the first large language model (LLM)-driven digital twin system for ophthalmology, which transforms longitudinal medical histories from 3,220 patients into structured clinical narratives. The system effectively handles irregularly sampled time-series data without requiring imputation, demonstrating enhanced modeling capacity under high clinical variability. In neovascular age-related macular degeneration (nAMD) prediction, it achieves a 6.0% reduction in mean absolute error (MAE) compared to existing baselines. For diabetic macular edema (DME), it outperforms Random Forest and XGBoost by 2.6% and 6.9% in MAE, respectively. These results validate the innovation and superiority of LLM-powered generative digital twins in longitudinal ophthalmic forecasting.

0 citationsRead paper

Communicating results in trials with multiple hypotheses or adaptive design features

May 05, 2026

This study addresses the underappreciated challenge of estimation and communication following multiplicity adjustment within the frequentist framework in complex clinical trials, where multiple endpoints, interim data looks, or group comparisons often introduce estimation bias and complicate interpretation, thereby undermining transparency in benefit–risk assessment. By integrating advanced methodologies such as adaptive designs and graphical approaches to multiple testing, the work illustrates through concrete examples the limitations of current strategies in conveying trial results meaningfully. The research underscores the need to critically reevaluate prevailing practices and foster interdisciplinary dialogue to enhance both the accuracy of effect estimation and the clarity of result communication, ultimately informing future methodological standards and regulatory guidance.

0 citationsRead paper

MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment

Apr 22, 2026

This work addresses the inherent conflicts among multiple alignment objectives—such as helpfulness, truthfulness, and harmlessness—in large language models, where conventional fixed scalarization methods often lead to systematic neglect of certain goals. The paper introduces, for the first time, a geometry-aware multi-objective optimization approach into the Direct Preference Optimization (DPO) framework, proposing a decoupled optimization method based on Multiple Gradient Descent Algorithm (MGDA). By dynamically identifying a shared descent direction across objectives, the method achieves a fair trade-off without requiring reinforcement learning or explicit reward models. Experiments on the UltraFeedback dataset demonstrate that the proposed approach attains state-of-the-art performance, achieving the highest win rates against golden responses both overall and on individual evaluation criteria.

0 citationsRead paper

Bayesian analysis of the causal reference-based model for missing data in clinical trials, accommodating partially observed post-intercurrent event data

Mar 27, 2026

This study addresses the challenge of partially missing data following intercurrent events in clinical trials, where conventional methods such as last observation carried forward yield large standard errors and convergence difficulties, while reference-based imputation relies on strong assumptions that may introduce bias. The authors propose the first extension of Bayesian causal models (BCM) to this setting, leveraging a fully Bayesian framework that integrates observed data with flexible prior information to achieve robust imputation. A key innovation is the introduction of an adjustable prior variance, which enhances estimation stability—particularly under data sparsity—and demonstrably outperforms existing approaches. Simulation studies show that, compared to traditional strategies, the proposed method substantially reduces standard errors and yields more stable estimates of treatment effects, especially when data are scarce.

0 citationsRead paper