AudioSpan: Spanning the Duration and Depth of Audio Comprehension

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
AudioSpan通过提供10分钟到2小时的音频与多层次认知问题,测试长音频理解能力,评估了12个大型音频语言模型。
📝 Abstract
General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal. Yet the benchmarks that test them still rely on clips of seconds, where scores saturate and models converge; recent long-form efforts extend duration but evaluate long audio much as short clips are. We introduce AudioSpan, a benchmark that spans both duration and depth: it pairs audio from 10 minutes to over 2 hours with 3,240 questions across three cognitive levels, namely perception, understanding, and reasoning. Two paths supply the questions, differing in how question content is sourced and how ground truth is obtained. Native QA extracts questions from the audio's content, posing each as a multiple-choice item and an open-ended one graded by detailed rubrics. Anchor QA instead injects ground truth, planting acoustic anchors into the audio and building a perception-to-reasoning chain scored only to the first error. A fully automated pipeline constructs every item through structured captioning, QA generation, and adversarial critic feedback. Evaluating 12 LALMs on AudioSpan, we find the hard part comes before reasoning: distilling a few relevant facts from a long, redundant signal. This difficulty grows with audio length and falls hardest on perception, especially temporal grounding. AudioSpan is available at https://huggingface.co/datasets/holvan/AudioSpan.
Problem

Research questions and friction points this paper is trying to address.

Audio Comprehension
Long-form Audio
Cognitive Levels
Innovation

Methods, ideas, or system contributions that make the work stand out.

AudioSpan
cognitive levels
long-form audio
automated pipeline
perception
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
W
Wen Huang
Qwen Team, Alibaba Group
Yunfei Chu
Yunfei Chu
Alibaba Group
machine learning
M
Meng Gao
Tsinghua University
H
Haolin He
The Chinese University of Hong Kong
Jin Xu
Jin Xu
Qwen Team, Alibaba Group
Multimodal InteractionLarge Language ModelSpeech SynthesisVideo/Audio Processing