CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration
This study addresses the fragmentation of specialties and limited modality support in existing medical foundation models by proposing a unified visual foundation model integrating pathology and radiology. Leveraging a multi-dimensional context attention mechanism, the framework unifies sparse and dense prediction tasks across 2D and high-dimensional inputs. Combined with multi-task joint training and parameter-efficient fine-tuning (PEFT), it enables federated learning on consumer-grade hardware. This work represents the first cross-specialty, multi-dimensional unified modeling approach for medical vision, achieving state-of-the-art performance across 12 benchmark datasets. Notably, fine-tuning less than 2.5% of parameters yields results comparable to full fine-tuning, while federated learning performance closely approximates centralized training, significantly enhancing adaptation efficiency in low-resource settings.