Institution profile

ARC Training Centre in Optimisation Technologies, Integrated Methodologies, and Applications (OPTIMA)

Industry researchaustralasia · au
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

RetFiner: A Vision-Language Refinement Scheme for Retinal Foundation Models

Jun 27, 2025

Existing OCT foundation models rely solely on image-based self-supervised learning, resulting in limited semantic understanding—particularly of anatomical structures and pathological concepts—and consequently suboptimal performance on complex downstream tasks; moreover, they require supervised fine-tuning for domain adaptation to specific patient populations. Method: We propose the first vision-language joint optimization framework tailored for retinal OCT, which integrates textual supervision into representation learning without additional annotations, via multi-objective collaborative training—including contrastive learning, image–text matching, and masked language modeling. Contribution/Results: The framework significantly enhances anatomical and pathological semantic comprehension, enabling zero-shot or lightweight direct adaptation. Evaluated on seven OCT classification tasks using RETFound, UrFound, and VisionFM, it achieves average linear probe accuracy gains of +5.8, +3.9, and +2.1 percentage points, respectively, consistently outperforming all baselines.

0 citationsRead paper
Recent publications

Latest Papers

RetFiner: A Vision-Language Refinement Scheme for Retinal Foundation Models

Jun 27, 2025

Existing OCT foundation models rely solely on image-based self-supervised learning, resulting in limited semantic understanding—particularly of anatomical structures and pathological concepts—and consequently suboptimal performance on complex downstream tasks; moreover, they require supervised fine-tuning for domain adaptation to specific patient populations. Method: We propose the first vision-language joint optimization framework tailored for retinal OCT, which integrates textual supervision into representation learning without additional annotations, via multi-objective collaborative training—including contrastive learning, image–text matching, and masked language modeling. Contribution/Results: The framework significantly enhances anatomical and pathological semantic comprehension, enabling zero-shot or lightweight direct adaptation. Evaluated on seven OCT classification tasks using RETFound, UrFound, and VisionFM, it achieves average linear probe accuracy gains of +5.8, +3.9, and +2.1 percentage points, respectively, consistently outperforming all baselines.

0 citationsRead paper