Understanding Task Aggregation for Generalizable Ultrasound Foundation Models

📅 2026-03-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing unified ultrasound foundation models often suffer performance degradation in multi-task joint training due to suboptimal task aggregation strategies, a problem exacerbated under limited data regimes. This work systematically investigates the feasibility of jointly learning heterogeneous tasks in ultrasound imaging and reveals, for the first time, that the efficacy of task aggregation critically depends on the interplay between training data scale and task type. Building upon the DINOv3 architecture, we propose M2DINO, a unified framework incorporating a task-conditioned Mixture-of-Experts module capable of supporting 27 diverse ultrasound tasks spanning segmentation, classification, detection, and regression. Empirical results demonstrate that training all tasks jointly yields more stable performance than clinically motivated groupings, with segmentation tasks exhibiting the highest susceptibility to negative transfer, while classification and regression tasks show greater robustness.

Technology Category

Application Category

📝 Abstract
Foundation models promise to unify multiple clinical tasks within a single framework, but recent ultrasound studies report that unified models can underperform task-specific baselines. We hypothesize that this degradation arises not from model capacity limitations, but from task aggregation strategies that ignore interactions between task heterogeneity and available training data scale. In this work, we systematically analyze when heterogeneous ultrasound tasks can be jointly learned without performance loss, establishing practical criteria for task aggregation in unified clinical imaging models. We introduce M2DINO, a multi-organ, multi-task framework built on DINOv3 with task-conditioned Mixture-of-Experts blocks for adaptive capacity allocation. We systematically evaluate 27 ultrasound tasks spanning segmentation, classification, detection, and regression under three paradigms: task-specific, clinically-grouped, and all-task unified training. Our results show that aggregation effectiveness depends strongly on training data scale. While clinically-grouped training can improve performance in data-rich settings, it may induce substantial negative transfer in low-data settings. In contrast, all-task unified training exhibits more consistent performance across clinical groups. We further observe that task sensitivity varies by task type in our experiments: segmentation shows the largest performance drops compared with regression and classification. These findings provide practical guidance for ultrasound foundation models, emphasizing that aggregation strategies should jointly consider training data availability and task characteristics rather than relying on clinical taxonomy alone.
Problem

Research questions and friction points this paper is trying to address.

task aggregation
ultrasound foundation models
task heterogeneity
training data scale
negative transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

task aggregation
ultrasound foundation model
Mixture-of-Experts
multi-task learning
negative transfer
🔎 Similar Papers
No similar papers found.
Fangyijie Wang
Fangyijie Wang
School of Medicine, University College Dublin
Computer VisionMedical ImagingDeep Learning
T
Tanya Akumu
Departament de Matemàtiques i Informàtica, Universitat de Barcelona, Barcelona, Spain
Vien Ngoc Dang
Vien Ngoc Dang
University of Barcelona
Computer VisionMedical ImagingDeep LearningFairness
A
Amelia Jimńez-Sánchez
Departament de Matemàtiques i Informàtica, Universitat de Barcelona, Barcelona, Spain
J
Jieyun Bai
Department of Cardiovascular Surgery, The First Affiliated Hospital of Jinan University, Jinan University, Guangzhou, China; Auckland Bioengineering Institute, University of Auckland, Auckland, New Zealand
G
Guénolé Silvestre
Research Ireland Centre for Research Training in Machine Learning; School of Computer Science, University College Dublin, Dublin, Ireland
Karim Lekadir
Karim Lekadir
ICREA Research Professor, Universitat de Barcelona
Biomedical data sciencehealthcare AItrustworthy AImedical image analysis
K
Kathleen M. Curran
Research Ireland Centre for Research Training in Machine Learning; School of Medicine, University College Dublin, Dublin, Ireland