MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出HEAL方法,通过干预多头输出和信息校准减少MLLMs的幻觉问题,提高模型可信度。
📝 Abstract
Multimodal Large Language Models (MLLMs) often struggle with hallucinations, thus hindering their reliable practical applications. Existing attention-based mitigation methods mainly rely on indirect signals (e.g., attention weights) that fail to accurately reflect the actual information shift underlying hallucination generation. In this paper, we propose HEAL, Head-lEvel information disentAnglement and caLibration for identifying and mitigating hallucinations. HEAL first employs causal noise intervention on multi-head outputs to filter out causally redundant heads. Subsequently, it disentangles information distribution within the remaining heads via the counterfactual Difference-in-Differences, categorizing heads into four types. Through analysis, we observe: hallucinations happen when information distribution drifts away from a healthy equilibrium in synergy heads, not strongly correlated with the quantity or strength of modality-specific heads. Motivated by this insight, HEAL injects dynamic information calibration factors into the value vectors of synergy heads, and actively regulates visual-language dependencies, steering the output distribution towards factual evidence. Extensive experiments demonstrate that HEAL effectively reduces hallucinations across multiple MLLMs, offering a simple and interpretable pathway to enhance model trustworthiness.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
hallucinations
information distribution drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Head-lEvel information disentAnglement and caLibration
causal noise intervention
counterfactual Difference-in-Differences
dynamic information calibration factors
🔎 Similar Papers
No similar papers found.
M
Meng'en Qin
Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology, Shenzhen, China
J
Junye Chen
Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology, Shenzhen, China
J
Jucheng Liu
Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology, Shenzhen, China
Y
Youlu Xing
Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology, Shenzhen, China
Song Wang
Song Wang
Professor of Computer Science and Engineering, University of South Carolina
Computer VisionImage ProcessingMachine Learning
Ruize Han
Ruize Han
SUAT
Computer VisionMultimedia AnalysisVideo UnderstandingActive Vision