Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexplored. Existing benchmarks conflate two distinct forms of evidence within a single modality: perceptual signals (e.g., a photograph or recording of a dog) and propositional signals (e.g., the declarative claim"this is a dog"), such that any measured modality bias is inherently confounded with evidence-form bias, precluding clean attribution to either source. To address this, we introduce Tri-PvP, an 8,000-sample tri-modal conflict benchmark crossing vision, audio, and text, where vision and audio each take perceptual or propositional form. Evaluating five OLLMs, we find robust visual bias across most models and evidence-type conditions. Crucially, we reveal a systematic asymmetry in evidence-form bias: models exhibit a stronger bias toward perceptual signal in vision but propositional in audio. Further analyses via layer-wise linear probing and contrastive decoding reveal that modality bias is already linearly decodable from early representation layers and can only be partially mitigated, calling for mitigation strategies beyond surface-level interventions.
Problem

Research questions and friction points this paper is trying to address.

modality bias
cross-modal conflict
perceptual signals
propositional signals
omni-modal large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tri-PvP
Modality Bias
Perceptual-Propositional Evidence Conflicts
Omni-Modal Large Language Models
💼 Related Jobs
No related jobs found.
Y
Yen-Ting Piao
National Taiwan University, Taipei, Taiwan
S
Shu-Yun Chen
National Taiwan University, Taipei, Taiwan
C
Chin-Hui Chu
National Taiwan University, Taipei, Taiwan
C
Chun-Wei Chen
National Taiwan University, Taipei, Taiwan
S
Shih-Yun Shan Kuan
National Taiwan University, Taipei, Taiwan
Hung-yi Lee
Hung-yi Lee
National Taiwan University
deep learningspoken language understandingspeech processing
Y
Yun-Nung Chen
National Taiwan University, Taipei, Taiwan