Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分析音频-视觉冲突测试AV-LLMs的组合泛化能力,发现模型在跨模态冲突下存在晚层优先主导问题,并通过机制可解释性分析定位了问题所在。
📝 Abstract
We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment configurations remain nearchance on the scored exact-string Yes/No subset of AVHBench, even though their output priors shift substantially. Similarly, off-the-shelf InternVideo2 experienced a 32.3% accuracy decrease specifically under cross-modal conflict, accompanied by a 17.3% instruction-following failure. We call this failure mode prior dominance: late-layer commitment to an internally preferred answer pattern that is weakly grounded in the conflicting inputs. To explain this behavior, we conduct a mechanistic interpretability analysis and find that commitment remains concentrated at 25.5 $\pm$ 1 layers. We show that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution. Code and data to reproduce our mechanistic audit and behavioral evaluations are available at https://github.com/AdarshSudheer09/AVHBench-dmai.
Problem

Research questions and friction points this paper is trying to address.

audio-visual conflict
compositional generalization
cross-modal conflict
Innovation

Methods, ideas, or system contributions that make the work stand out.

prior dominance
cross-modal conflict
mechanistic interpretability analysis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Adarsh Sudheer
Independent Researcher
D
David Li
Independent Researcher
O
Omar Elbanna
Independent Researcher
I
Ishaan Kodarapu
Independent Researcher
A
Arjun Bahuguna
Independent Researcher
Vasu Sharma
Vasu Sharma
Facebook AI Research (FAIR)
Generative AILLMsComputer VisionNatural Language ProcessingMultimodal ML