Task-disentangled Low-Rank Adaptation for Versatile Audio-visual Multi-modal Learning Tasks within a Unified Framework

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究提出了一种任务解耦的低秩适应机制,通过统一框架处理多种音频-视觉多模态学习任务,利用大语言模型的能力实现任务间的有效协作。
📝 Abstract
Inspired by human multi-modal perception, Audio-Visual Multi-Modal Learning (AVMML) integrates auditory and visual information to leverage complementary cross-modal cues, enabling more robust and comprehensive scene perception. Existing studies predominantly tackle each AVMML task in isolation, which stands in stark contrast to humans' unified cognitive capacity for handling versatile perception. However, naive joint training across multiple AVMML tasks often suffers from mutual interference, arising from the intricate inter-task relationships. To address this, we propose a unified framework that simultaneously accommodates versatile AVMML tasks. Specifically, benefiting from powerful representation and generalization capabilities of large language models, we design a task-disentangled Low-Rank Adaptation (LoRA) mechanism that enables dynamic integration of both task-specific and task-shared knowledge, thereby facilitating effective multi-task collaboration. The proposed task-disentangled LoRA comprises three components: a task-general low-rank matrix, task-specific modulation matrices, and cross-task collaboration experts, which respectively capture universal audio-visual knowledge, decouple task-specific pattern, and exploit inherent inter-task correlations. By unifying explicit collaboration from both model and task perspectives, our approach not only surpasses existing unified audio-visual models across multiple AVMML tasks, but also outperforms most task-specific models on certain AVMML tasks.
Problem

Research questions and friction points this paper is trying to address.

Audio-Visual Multi-Modal Learning
task interference
unified framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-disentangled Low-Rank Adaptation (LoRA)
multi-modal learning
unified framework
cross-task collaboration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hanyu Xuan
School of Big Data and Statistics, Anhui University, Hefei 230039, China
M
Mengqi Zhang
School of Big Data and Statistics, Anhui University, Hefei 230039, China
J
Junjun Mao
School of Big Data and Statistics, Anhui University, Hefei 230039, China
F
Fei Wang
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, Hefei, 230026, China
K
Kun Li
College of Information Technology, United Arab Emirates University, Abu Dhabi, 15551, United Arab Emirates
G
Guanghui Yue
School of Biomedical Engineering, Shenzhen University, Shenzhen 518060, China
Zhiliang Wu
Zhiliang Wu
Research Scientist, Siemens Technology
Representation learningMachine learningGaussian ProcessesHealthcare
Hehe Fan
Hehe Fan
Zhejiang University
Deep learningComputer visionMultimediaAI for science