Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs

📅 2026-02-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the “modality collapse” phenomenon in multimodal large language models, wherein decoders trained solely toward textual alignment fail to effectively leverage modality-specific cues—such as speaker identity and emotion in speech or visual textures in images. The study formalizes this issue through an information-theoretic lens as a mismatch between the decoder and the input distribution, and introduces Generalized Mutual Information (GMI) as a theoretical upper bound on accessible information. Through linear probing, LoRA fine-tuning, and controlled contrastive experiments across five audio-visual models, the authors demonstrate that optimizing the decoding objective significantly enhances the extractability of non-textual information: LoRA-based intervention boosts emotion information accessibility by 7.5% without degrading other attributes, thereby establishing that the training objective fundamentally governs information usability.

Technology Category

Application Category

📝 Abstract
Multimodal LLMs can process speech and images, but they cannot hear a speaker's voice or see an object's texture. We show this is not a failure of encoding: speaker identity, emotion, and visual attributes survive through every LLM layer (3--55$\times$ above chance in linear probes), yet removing 64--71% of modality-specific variance improves decoder loss. The decoder has no learned use for these directions; their presence is noise. We formalize this as a mismatched decoder problem: a decoder trained on text can only extract information along text-aligned directions. Accessible information is bounded by the Generalized Mutual Information (GMI), with degradation scaling with distributional distance and decoder sensitivity. The bound is a property of the decoder's scoring rule, not of any particular architecture; it applies whether non-text inputs arrive through a learned projection, a discrete codebook, or no explicit adapter at all. We validate this across five models spanning speech and vision. A controlled experiment (two Prismatic VLMs differing only in encoder text-alignment) confirms the bottleneck is the decoder's scoring rule, not the encoder or projection. A LoRA intervention demonstrates the fix: training with an emotion objective improves emotion accessibility ($+$7.5%) without affecting other attributes, confirming that the training objective determines what becomes accessible.
Problem

Research questions and friction points this paper is trying to address.

Modality Collapse
Mismatched Decoding
Multimodal LLMs
Generalized Mutual Information
Decoder Bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

mismatched decoding
Generalized Mutual Information
multimodal LLMs
scoring rule
modality collapse
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jayadev Billa
Unaffiliated researcher; previously at ISI@USC, Yahoo, Nuance, and BBN