Visual Instruction Tuning Aligns Modalities through Abstraction
This work investigates how visual instruction tuning integrates image information into the hierarchical architecture of large language models—a mechanism not well understood in prior studies. Through a systematic analysis employing probing, causal intervention, representational geometry comparison, and selective layer fine-tuning, the study reveals that visual features are predominantly embedded in the model’s intermediate semantic layers in a localized manner. Building on this insight, the authors demonstrate that fine-tuning only these intermediate layers achieves performance comparable to full-model fine-tuning across multiple vision-centric benchmarks, while substantially reducing training costs. These findings underscore the pivotal role of intermediate layers in cross-modal alignment and offer an efficient strategy for multimodal adaptation.