Institution profile

University Hospital Tübingen

Academic institutioneurope · de
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration

Aug 17, 2026

This study addresses the fragmentation of specialties and limited modality support in existing medical foundation models by proposing a unified visual foundation model integrating pathology and radiology. Leveraging a multi-dimensional context attention mechanism, the framework unifies sparse and dense prediction tasks across 2D and high-dimensional inputs. Combined with multi-task joint training and parameter-efficient fine-tuning (PEFT), it enables federated learning on consumer-grade hardware. This work represents the first cross-specialty, multi-dimensional unified modeling approach for medical vision, achieving state-of-the-art performance across 12 benchmark datasets. Notably, fine-tuning less than 2.5% of parameters yields results comparable to full fine-tuning, while federated learning performance closely approximates centralized training, significantly enhancing adaptation efficiency in low-resource settings.

0 citationsRead paper

Register and CLS tokens yield a decoupling of local and global features in large ViTs

May 09, 2025

Large Vision Transformers (ViTs) suffer from distorted attention maps, reduced interpretability, and suboptimal dense prediction performance due to an inherent trade-off between global and local feature modeling—specifically, the suppression of local information fusion by dominant global tokens. Method: Through systematic attention analysis, token attribution, and architectural diagnosis—empirically validated on models including DINOv2—we identify a fundamental structural equivalence between the CLS token and explicit register tokens. Contribution/Results: We establish, for the first time, that the CLS token functions as an *implicit register*, jointly governing global representation with explicit registers while impeding local feature integration. This challenges prevailing assumptions about attention interpretability in ViTs and provides a novel theoretical framework for understanding representational bottlenecks. Our findings inform principled design strategies for next-generation vision models that simultaneously achieve robust local–global modeling and enhanced interpretability.

0 citationsRead paper
Recent publications

Latest Papers

CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration

Aug 17, 2026

This study addresses the fragmentation of specialties and limited modality support in existing medical foundation models by proposing a unified visual foundation model integrating pathology and radiology. Leveraging a multi-dimensional context attention mechanism, the framework unifies sparse and dense prediction tasks across 2D and high-dimensional inputs. Combined with multi-task joint training and parameter-efficient fine-tuning (PEFT), it enables federated learning on consumer-grade hardware. This work represents the first cross-specialty, multi-dimensional unified modeling approach for medical vision, achieving state-of-the-art performance across 12 benchmark datasets. Notably, fine-tuning less than 2.5% of parameters yields results comparable to full fine-tuning, while federated learning performance closely approximates centralized training, significantly enhancing adaptation efficiency in low-resource settings.

0 citationsRead paper

Register and CLS tokens yield a decoupling of local and global features in large ViTs

May 09, 2025

Large Vision Transformers (ViTs) suffer from distorted attention maps, reduced interpretability, and suboptimal dense prediction performance due to an inherent trade-off between global and local feature modeling—specifically, the suppression of local information fusion by dominant global tokens. Method: Through systematic attention analysis, token attribution, and architectural diagnosis—empirically validated on models including DINOv2—we identify a fundamental structural equivalence between the CLS token and explicit register tokens. Contribution/Results: We establish, for the first time, that the CLS token functions as an *implicit register*, jointly governing global representation with explicit registers while impeding local feature integration. This challenges prevailing assumptions about attention interpretability in ViTs and provides a novel theoretical framework for understanding representational bottlenecks. Our findings inform principled design strategies for next-generation vision models that simultaneously achieve robust local–global modeling and enhanced interpretability.

0 citationsRead paper