Beyond Accuracy: ARIA-Rubrics for Evaluating Audio Reasoning in Large Audio Language Models

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型音频语言模型的推理能力评估问题,提出ARIA-Rubrics框架,通过六个互补指标自动透明地评估音频推理质量。
📝 Abstract
Large Audio Language Models (LALMs) have shown strong performance on audio reasoning benchmarks, but accuracy alone cannot distinguish true reasoning from superficial pattern matching, often overestimating reasoning ability since high scores may result from guessing rather than genuine audio understanding. Evaluating the reasoning process itself is essential for improving LALMs' reasoning ability, yet remains challenging. Existing methods either rely on costly human annotation or opaque LLM-as-judge approaches, making them impractical, biased, and lacking transparency. Moreover, audio reasoning introduces unique challenges absent in text-based settings, perceptual hallucination and cross-modal alignment between audio understanding and textual inference, hence text-based evaluation frameworks cannot be directly applied. Therefore, we propose ARIA-Rubrics (Audio Reasoning Integrity Assessment), a lightweight, annotation-free gold reasoning chains, automatic and transparent framework comprising six complementary metrics that evaluate audio reasoning quality across perceptual grounding, reasoning coherence, and answer consistency. We use Chain-of-Thought prompting as an externalization mechanism to make the reasoning process observable. Experiments on 9 models across 2 benchmarks identify three reasoning modes of current LALMs with actionable directions for future development, with ARIA-Rubrics achieving high correlation with human judgments. The code is available at the Github Repository.
Problem

Research questions and friction points this paper is trying to address.

Large Audio Language Models
Audio Reasoning
Evaluation Metrics
Perceptual Hallucination
Cross-modal Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

ARIA-Rubrics
Audio Reasoning
Chain-of-Thought
Perceptual Grounding
Reasoning Coherence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yupei Li
Imperial College London, UK
Qiyang Sun
Qiyang Sun
Imperial College London
M
Mohamed Mady
Technische Universität München, München, Germany
C
Chenxi Wang
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, AE
Z
Zhengwei Gong
Shanghai Jiao Tong University, Shanghai, CN
Berrak Sisman
Berrak Sisman
Assistant Professor (ECE & DSAI), Johns Hopkins University
Machine LearningAffective ComputingSpeech SynthesisVoice ConversionAnti-spoofing
Björn Schuller
Björn Schuller
Professor, Technische Universität München (TUM) / Imperial College London & CSO, audEERING
Health InformaticsDigital HealthAIAffective ComputingComputer Audition