Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过提出SHS-Index量化视觉Transformer注意力头的语义专业化,设计了Ariadne Attention以减少计算同时保持性能。
📝 Abstract
Hybrid attention dominates frontier LLMs, yet Vision Transformers (ViTs) in multimodal LLMs lack a satisfactory hybrid design, with no consensus on why certain attention patterns work better. To fill this gap, we study ViT attention heads and find they differentiate into object- and background-specialist roles, a pattern most pronounced under full attention; we call this Semantic Head Specialization (SHS). We propose SHS-Index to quantify this specialization, show that it distinguishes full-attention from chunk-window ViTs, and find that it strongly tracks downstream benchmark performance. We then identify three structural factors that shape SHS---window interaction, token serialization, and local softmax allocation---and use them as design principles for hybrid attention. Guided by these factors, we design Ariadne Attention, a hybrid that matches full attention on 22 image and video tasks at 6.5x less attention compute. Our findings establish head specialization as a measurable property for diagnosing and designing principled hybrid ViT attention at the multimodal-LLM scale.
Problem

Research questions and friction points this paper is trying to address.

Hybrid Attention
Vision Transformers
Multimodal LLMs
Semantic Head Specialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Head Specialization
SHS-Index
Ariadne Attention
Hybrid Attention
🔎 Similar Papers
No similar papers found.
C
Chenhong He
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; Xiaomi Corporation, LLM-Core Team
L
Lei Li
Department of Computer Science, The University of Hong Kong; Xiaomi Corporation, LLM-Core Team
S
Shicheng Li
Xiaomi Corporation, LLM-Core Team
H
Hanglong Lv
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; Xiaomi Corporation, LLM-Core Team
Lingpeng Kong
Lingpeng Kong
Google DeepMind, The University of Hong Kong
Natural Language ProcessingMachine Learning
Q
Qi Liu
Department of Computer Science, The University of Hong Kong
Tong Yang
Tong Yang
Peking University, Beijing, China. PKU. 北京大学
SketchNetwork measurementBloom filterIP lookupHash Table
Shuhuai Ren
Shuhuai Ren
Peking University
Deep LearningNatural Language Processing