LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态大语言模型在视觉指令调优中的灾难性遗忘问题,提出LLaVAFlow框架,通过信息压缩轨迹保持跨模态对齐流,提高下游任务性能。
📝 Abstract
While Multimodal Large Language Models (MLLMs) exhibit strong generalization, visual instruction tuning for downstream tasks inevitably causes catastrophic forgetting, impairing overall generalization. While existing methods regulate weight updates to reduce forgetting, they overlook the fundamental cross-modal alignment in MLLMs. Based on prior work and our observations, we argue that cross-modal alignment is implicitly captured in the information-compression trajectory. To preserve the alignment flow embedded in the trajectory, we propose LLaVAFlow, an information-theoretic distillation framework. First, we compress the mutual information between the extracted relations and MLLM embeddings, encouraging a learnable module to produce a refined alignment flow that benefits downstream tasks. Second, we maximize the mutual information between the extracted alignment flows of the pretrained and fine-tuned MLLMs, enabling the transfer of compact alignment information. Extensive experiments show that LLaVAFlow is an effective plug-and-play framework that preserves alignment flow and enhances both downstream performance and generalization.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
catastrophic forgetting
cross-modal alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal alignment
information-theoretic distillation
mutual information
alignment flow
🔎 Similar Papers
2024-08-29arXiv.orgCitations: 7
💼 Related Jobs
No related jobs found.
M
Muyao Yuan
MOEKLINNS, and School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an, China
M
Muyan Jiao
MOEKLINNS, and School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an, China
J
Jiangyong Ying
E-surfing Vision Technology Co., Ltd, China Telecom, Hangzhou, China
Weizhan Zhang
Weizhan Zhang
Professor,Department of Computer Science and Technology, Xi'an Jiaotong University
Multimedia networking
Y
Yuanhong Zhang
MOEKLINNS, and School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an, China
L
Lan Ma
China Telecom, Xi’an, China
Yuan Gao
Yuan Gao
University of Science and Technology of China
Graph MiningAnomaly DetectionOut-of-distribution Generalization
H
Haipeng Du
MOEKLINNS, and School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an, China