Auditory Illusion Benchmark for Large Audio Language Models

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过创建首个针对大型音频语言模型的听觉错觉基准AIB,使用包含音乐、声音和语音的十种代表性错觉来测试模型是否复制人类感知倾向,揭示了现有模型在模仿人类听觉认知上的局限。
📝 Abstract
Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB, the first auditory illusion benchmark for LALMs, covering ten representative illusions across music, sound, and speech, each annotated for the presence of knowledge-based priors. Our methodology pairs model evaluation with controlled human listening studies, enabling direct comparison of responses. Results show systematic differences: while most LALMs remain signal-faithful on low-level acoustic illusions, several exhibit more human-like responses when linguistic or musical priors are involved, although no model matches the human perceptual profile. These findings highlight the current limitations of LALMs as cognitive models. By establishing auditory illusions as a rigorous testbed, our work offers a new perspective for probing neural black-box models and advancing understanding of auditory cognition. AIB is publicly available at https://github.com/gillosae/aib.
Problem

Research questions and friction points this paper is trying to address.

Auditory Illusions
Large Audio Language Models
Perceptual Tendencies
Benchmark
Human Cognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Auditory Illusion Benchmark
Large Audio Language Models
Perceptual Tendencies
Knowledge-based Priors
Cognitive Models
🔎 Similar Papers