Institution profile

International Laboratory on Learning Systems

Academic institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic Segmentation

May 28, 2025

This work pioneers test-time adaptation (TTA) for open-vocabulary semantic segmentation (OVSS), addressing the previously unexplored challenge of unsupervised TTA in dense prediction with vision-language models (VLMs). We propose Multi-Level Multi-Prompt entropy Minimization (MLMP), a plug-and-play, training-free, and label-free method that jointly optimizes CLIP’s global text–image alignment and pixel-level visual representations, while integrating intermediate-layer features and diverse textual prompts. To enable systematic evaluation, we establish the first OVSS TTA benchmark—comprising seven datasets, fifteen image corruptions, and eighty-two distribution shift scenarios. Under unified evaluation, single-sample TTA consistently improves mean Intersection-over-Union (mIoU) by 2.1–4.7 percentage points, substantially outperforming existing image-classification TTA methods. These results demonstrate MLMP’s strong cross-distribution robustness and generalization capability in dense prediction settings.

0 citationsRead paper
Recent publications

Latest Papers

Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic Segmentation

May 28, 2025

This work pioneers test-time adaptation (TTA) for open-vocabulary semantic segmentation (OVSS), addressing the previously unexplored challenge of unsupervised TTA in dense prediction with vision-language models (VLMs). We propose Multi-Level Multi-Prompt entropy Minimization (MLMP), a plug-and-play, training-free, and label-free method that jointly optimizes CLIP’s global text–image alignment and pixel-level visual representations, while integrating intermediate-layer features and diverse textual prompts. To enable systematic evaluation, we establish the first OVSS TTA benchmark—comprising seven datasets, fifteen image corruptions, and eighty-two distribution shift scenarios. Under unified evaluation, single-sample TTA consistently improves mean Intersection-over-Union (mIoU) by 2.1–4.7 percentage points, substantially outperforming existing image-classification TTA methods. These results demonstrate MLMP’s strong cross-distribution robustness and generalization capability in dense prediction settings.

0 citationsRead paper