Mind the Heads: Topological Representation Alignment for Multimodal LLMs
Existing multimodal large language models typically perform coarse-grained representation alignment only at fixed linguistic layers, overlooking the fine-grained structure within individual Transformer attention heads, which limits cross-modal alignment efficacy. This work proposes HeRA, the first method to achieve cross-modal topological alignment at the level of individual attention heads. Grounded in the Bregman representation hypothesis, HeRA enhances alignment quality by preserving local neighborhood relationships across modalities. It introduces a differentiable Mutual K-Nearest Neighbor (MKNN) contrastive objective that dynamically identifies and optimizes critical attention heads, revealing that the least-aligned heads yield the greatest performance gains. Experiments demonstrate that HeRA significantly improves visual-centric task performance across multiple state-of-the-art multimodal large language models and 18 benchmarks, while effectively mitigating visual hallucination and overreliance on linguistic priors.