Self-supervised Pre-training Helps Retinal Disease Progression Modelling Most When Data Is Scarce

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过自监督预训练方法解决视网膜疾病进展建模中纵向数据稀缺问题,特别是在少量标记数据下使用冻结编码器效果最佳。
📝 Abstract
Modelling how a disease progresses over time requires longitudinal imaging cohorts, which are scarce and small, whereas cross-sectional data -- one image per participant -- is abundant. Self-supervised pre-training on such data offers a way to bridge this gap, but it is unclear which strategy best supports progression modelling, or how that answer depends on the amount of labelled longitudinal data. We study this for age-related macular degeneration (AMD), pre-training encoders on the large cross-sectional NAKO cohort and predicting time to late AMD on the longitudinal AREDS dataset. We compare in-house self-supervised encoders against a general-purpose (DINOv2) and a domain-specific (RETFound) foundation model, across contrastive, masked-autoencoding, and self-distillation objectives, under frozen and fine-tuned protocols, and across labelled training sets from 100 to 32,250 examples. Which model performs best depends on how the encoder is used. When the encoder is frozen and labels are few -- the regime typical of longitudinal cohorts -- pre-trained representations reach clinically reasonable discrimination from a few hundred labelled samples, while models trained from scratch do not; this advantage fades under fine-tuning. Transfer is governed by the self-supervision objective rather than corpus scale or domain match, so that an encoder pre-trained on a modest cross-sectional cohort matches or exceeds a far larger in-domain foundation model. Together, these results offer a practical recipe for building progression models where longitudinal data is scarce: a frozen self-supervised encoder with a lightweight survival head.
Problem

Research questions and friction points this paper is trying to address.

self-supervised pre-training
retinal disease progression
longitudinal data scarcity
cross-sectional data
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-supervised pre-training
disease progression modelling
longitudinal data scarcity
frozen encoder
contrastive learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
I
Ifeoma Veronica Nwabufo
Hertie Institute for AI in Brain Health, Faculty of Medicine, University of Tübingen, Germany; Tübingen AI Center, University of Tübingen, Germany
J
Julius Gervelmeyer
Hertie Institute for AI in Brain Health, Faculty of Medicine, University of Tübingen, Germany; Tübingen AI Center, University of Tübingen, Germany
Sarah Müller
Sarah Müller
University of Tübingen
Philipp Berens
Philipp Berens
Hertie Institute for AI in Brain Health, University of Tübingen
Computational NeuroscienceData ScienceMachine LearningDigital MedicineMedical AI