TAD: Token-Adaptive Contrastive Decoding with Confidence-Guided Gating for Hallucination Mitigation in Large Audio-Language Models

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Token-Adaptive Decoding (TAD)方法,通过对比实际音频和静音参考下的logits来减少大型音频语言模型中的幻觉问题。
📝 Abstract
Large audio-language models (LALMs) can hallucinate audio objects, answering"yes"to absent sound events, thus undermining reliability in audio question answering. We propose Token-Adaptive Decoding (TAD), a training-free strategy for hallucination mitigation that grounds the initial yes/no decision by contrasting logits under real audio with a matched silent reference. TAD introduces a token-adaptive, confidence-guided gate that is decision-critical at the first decoding step and class-conditional on affirmative tokens, using the audio-silent margin to avoid overcorrection when evidence is weak or already sufficient. Experiments on AudioCaps-Hallucination show that, relative to Audio-Aware Decoding (AAD), a contrastive baseline with fixed contrast strength, TAD improves F1 for Qwen2 by 0.059 to 0.117 across Popular, Adversarial, and Random splits, and for Gemma by 0.025 to 0.064, while on Clotho-AQA it raises F1 from 0.810 to 0.816 on Qwen2 and remains comparable to AAD on Gemma.
Problem

Research questions and friction points this paper is trying to address.

hallucination
audio question answering
large audio-language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-Adaptive Decoding
Confidence-Guided Gating
Hallucination Mitigation
Large Audio-Language Models
Contrastive Decoding
H
Heyu Chang
Information Engineering University, China
N
Nianwen Si
Information Engineering University, China
Hao Zhang
Hao Zhang
Zachry Department of Civil and Environmental Engineering, Texas A&M University
Digital twinTraffic safetyAI in traffic engineering
W
Wenlin Zhang
Information Engineering University, China
D
Dan Qu
Information Engineering University, China