Hyper-LLaVA: Hyperbolic Uncertainty-aware Modality-Balanced Routing for Multimodal Continual Instruction Tuning

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态持续指令调优中的参数路由问题,提出Hyper-LLaVA方法,通过双曲空间中样本-任务分布相似性及模态间平衡来优化任务匹配。
📝 Abstract
Multimodal Continual Instruction Tuning (MCIT) aims to exploit the incrementally accumulated knowledge to process multimodal inputs of diverse tasks, where parameter routing plays an important role. State-of-the-art methods rely on sample-to-task center similarity and cross-modal fusion with equal weight during routing. However, such solutions face two fundamental flaws: (1) Within each modality, the sample-to-task center distance is sub-optimal for routing since the abundant intra-task diversity information is underleveraged. (2) Different modalities exhibit varying reliability across tasks, where the modality with inter-task ambiguity can easily misguide the routing result. To address these problems, we propose Hyperbolic Uncertainty-aware Modality-Balanced Routing (Hyper-LLaVA) to improve parameter routing capacity based on cross-modality task feature uncertainty modeling. Specifically, to improve intra-modality task matching, Hyper-LLaVA accesses the sample-to-task distribution similarity in the Hyperbolic space. Besides, to alleviate the degradation brought by unreliable modalities, Hyper-LLaVA quantifies the task matching ambiguity within each modality to achieve adaptive balancing between task matching across modalities. Based on the complementary intra- and inter-modality task matching enhancement, our Hyper-LLaVA outperforms state-of-the-art approaches by large margins. Our source code is available at https://github.com/zhoujiahuan1991/ICML2026-Hyper-LLaVA
Problem

Research questions and friction points this paper is trying to address.

Multimodal Continual Instruction Tuning
parameter routing
intra-task diversity
modality reliability
task matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hyperbolic Uncertainty-aware Modality-Balanced Routing
Multimodal Continual Instruction Tuning
Cross-modality Task Feature Uncertainty Modeling
Sample-to-task Distribution Similarity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kunlun Xu
Wangxuan Institute of Computer Technology, Peking University, Beijing, China
Y
Yanqin Zhang
Wangxuan Institute of Computer Technology, Peking University, Beijing, China
Wenwen Qiang
Wenwen Qiang
Institute of Software, Chinese Academy of Sciences
Artificial IntelligenceMachine LearningCausal InferenceLLM/MLLM
Jiahuan Zhou
Jiahuan Zhou
Peking University
Computer VisionMachine LearningDeep Learning