When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文针对文本误导问题,提出了一种基于音频的对话理解方法,通过构建Audio Twin框架来整合声学证据,从而提高模型在冲突情况下的准确性。
📝 Abstract
Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or single-modality solutions, obscuring whether models genuinely ground predictions in speech. We formalize this failure mode as cross-modal disagreement, where transcripts suggest plausible but incorrect surface interpretations while acoustic cues such as prosody or speaking style support different answers. We develop a scalable framework that identifies text-biased surface interpretations and converts disagreement regions into conflict QA examples. We also include consistent cases where transcript-based and speech-grounded interpretations agree, enabling evaluation beyond adversarial audio dependence. This results in ContraTalk, a controlled benchmark containing 501 questions across five discourse dimensions: interaction behavior, emotion state, dialogue act, social stance, and conversational intent. We further develop an agentic-style reasoning framework that converts speech into an Audio Twin, a text-readable representation of localized acoustic cues that exposes acoustic evidence to the reasoning model. Experiments show that strong text-only LLMs exceed 90% accuracy in consistent cases but drop to 33-48% in conflict cases. Direct AudioLLMs provide only partial grounding, still selecting the transcript-biased trap in roughly 30-40% of conflict cases. Our Audio Twin framework improves conflict-case accuracy while reducing trap selection, but its consistent-case behavior remains backbone-dependent. These results identify transcript-based shortcuts as an important failure mode in spoken dialogue understanding and show that explicit acoustic evidence aggregation provides a more controllable interface for diagnosing and improving speech-grounded reasoning.
Problem

Research questions and friction points this paper is trying to address.

cross-modal disagreement
audio-grounded dialogue
transcript-based shortcuts
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal disagreement
Audio Twin
acoustic evidence aggregation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yen-Ju Lu
Center for Language and Speech Processing, Johns Hopkins University
Y
Yuzhe Wang
Center for Language and Speech Processing, Johns Hopkins University
Y
Yaohan Guan
Center for Language and Speech Processing, Johns Hopkins University
X
Xiluo He
Center for Language and Speech Processing, Johns Hopkins University
Jiarui Hai
Jiarui Hai
Johns Hopkins University
computer auditiongenerative modelsmusic information retrieval
M
Mingrui Liang
Center for Language and Speech Processing, Johns Hopkins University
K
Kaavya Chaparala
Center for Language and Speech Processing, Johns Hopkins University
Thomas Thebaud
Thomas Thebaud
Assistant Research Scientist, ECE Dept., Johns Hopkins University, Baltimore
Adversarial and Backdoor attacksSpeech Emotion RecognitionAudio LLMsSpeaker Characterisation
L
Laureano Moro-Velazquez
Center for Language and Speech Processing, Johns Hopkins University
Najim Dehak
Najim Dehak
Associate Professor at ECE department, Johns Hopkins University.
Machine learningspeech processingspeaker recognitionlanguage recognitionemotion recognition
J
Jesus Villalba
Center for Language and Speech Processing, Johns Hopkins University