A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

๐Ÿ“… 2026-08-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บ่ง†่ง‰ไพ่ต–ๆ„Ÿ็Ÿฅๆก†ๆžถ๏ผŒ้€š่ฟ‡่ง†่ง‰็บฆๆŸๆœ€ไผ˜ไผ ่พ“ๅ’Œ่ง†่ง‰่ฐƒ่Š‚้€‚ๅบ”ๆ–นๆณ•่งฃๅ†ณๅคšๆจกๆ€ๆ— ็›‘็ฃๆŒ็ปญๅŽ่ฎญ็ปƒไธญ็š„่ทจๆจกๆ€็พ้šพๆ€ง้—ๅฟ˜้—ฎ้ข˜ใ€‚
๐Ÿ“ Abstract
In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Existing unsupervised post-training methods for MLLMs typically optimize target tokens uniformly, overlooking their heterogeneous visual dependence (VD). However, we reveal that token-level VD is crucial for MU-CPT. Specifically, its structural distortion serves as an indicator of cross-modal catastrophic forgetting, and its inherent heterogeneity acts as a compass to guide new-task learning. Leveraging this property, we propose a Visual Dependence-Aware (VDA) framework with two main components. First, Visually Constrained Optimal Transport (VC-OT) formulates the VD structural distortion of old-task VD during new-task learning as an optimal transport problem to mitigate cross-modal forgetting. By designing a region-aware ground cost and a dependence-stratified transport penalty, it prevents global shifts in visual focus while strictly prohibiting visual reliance from degenerating into language bias. Second, Visually Modulated Adaptation (VMA) exploits VD heterogeneity to emphasize visually grounded new-task learning, promoting new-task plasticity. Together, our method simultaneously maintains old-task stability and new-task plasticity during challenging MU-CPT. Extensive experiments under our MU-CPT setting validate the effectiveness of VDA.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Unsupervised Continual Post-Training
Visual Dependence
Cross-modal Catastrophic Forgetting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Dependence-Aware
Multimodal Unsupervised Continual Post-Training
Visually Constrained Optimal Transport
Visually Modulated Adaptation
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
K
Kaichen Li
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; Pengcheng Laboratory, Shenzhen, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Z
Zhilin Zhu
Harbin Institute of Technology, Shenzhen, China
J
Jianhao Huang
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; Pengcheng Laboratory, Shenzhen, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Z
Zhengqin Lai
Harbin Institute of Technology, Shenzhen, China
Baochen Xiong
Baochen Xiong
Institute of Automation, Chinese Academy of Sciences, Peng Cheng Lab
Federated LearningMultimedia
Z
Zibo Shao
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; Pengcheng Laboratory, Shenzhen, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Yaguang Song
Yaguang Song
Peng Cheng Laboratory
Deep LearningMulti-Modal Pre-training
L
Linhui Xiao
Pengcheng Laboratory, Shenzhen, China
X
Xiaoshan Yang
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; Pengcheng Laboratory, Shenzhen, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Changsheng Xu
Changsheng Xu
Professor, Institute of Automation, Chinese Academy of Sciences
MultimediaComputer vision