Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages
研究通过激活修补和注意力分析,探讨了多语言大模型在处理主谓一致时的共享机制,发现有显性人称/数屈折变化的语言表现出更相似的处理电路。
研究通过激活修补和注意力分析,探讨了多语言大模型在处理主谓一致时的共享机制,发现有显性人称/数屈折变化的语言表现出更相似的处理电路。
This work pioneers test-time adaptation (TTA) for open-vocabulary semantic segmentation (OVSS), addressing the previously unexplored challenge of unsupervised TTA in dense prediction with vision-language models (VLMs). We propose Multi-Level Multi-Prompt entropy Minimization (MLMP), a plug-and-play, training-free, and label-free method that jointly optimizes CLIP’s global text–image alignment and pixel-level visual representations, while integrating intermediate-layer features and diverse textual prompts. To enable systematic evaluation, we establish the first OVSS TTA benchmark—comprising seven datasets, fifteen image corruptions, and eighty-two distribution shift scenarios. Under unified evaluation, single-sample TTA consistently improves mean Intersection-over-Union (mIoU) by 2.1–4.7 percentage points, substantially outperforming existing image-classification TTA methods. These results demonstrate MLMP’s strong cross-distribution robustness and generalization capability in dense prediction settings.
研究通过激活修补和注意力分析,探讨了多语言大模型在处理主谓一致时的共享机制,发现有显性人称/数屈折变化的语言表现出更相似的处理电路。
This work pioneers test-time adaptation (TTA) for open-vocabulary semantic segmentation (OVSS), addressing the previously unexplored challenge of unsupervised TTA in dense prediction with vision-language models (VLMs). We propose Multi-Level Multi-Prompt entropy Minimization (MLMP), a plug-and-play, training-free, and label-free method that jointly optimizes CLIP’s global text–image alignment and pixel-level visual representations, while integrating intermediate-layer features and diverse textual prompts. To enable systematic evaluation, we establish the first OVSS TTA benchmark—comprising seven datasets, fifteen image corruptions, and eighty-two distribution shift scenarios. Under unified evaluation, single-sample TTA consistently improves mean Intersection-over-Union (mIoU) by 2.1–4.7 percentage points, substantially outperforming existing image-classification TTA methods. These results demonstrate MLMP’s strong cross-distribution robustness and generalization capability in dense prediction settings.