Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses critical limitations in existing audio-visual social understanding benchmarks, which suffer from high noise levels and poorly designed questions, while complex reasoning approaches often incur substantial costs with marginal gains. The authors systematically evaluate the reasoning capabilities of multimodal large language models on social audio-visual question answering, introducing IntentBench-Prime—a high-quality benchmark constructed through rigorous data cleaning—and comparing diverse training strategies. Their findings reveal that a simple vanilla supervised fine-tuning (SFT) baseline matches or surpasses state-of-the-art complex methods across three benchmarks. Notably, using only textual captions achieves performance comparable to full video inputs, suggesting that linguistic modalities encode strong social priors. The study further proposes a cost-effective evaluation paradigm and publicly releases the denoised IntentBench-Prime benchmark to support future research.
📝 Abstract
Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become the dominant approach, with HumanOmniV2 and its IntentBench benchmark as a prominent reference point. In this context, we report three findings. First, IntentBench is highly noisy: $\sim$7% of questions are broken and $\sim$23% are trivially answerable without the video input. We remove the affected questions and release Intentbench-Prime. Second, current reasoning approaches are expensive and surprisingly ineffective. A simple Vanilla SFT baseline matches or outperforms existing reasoning methods across three benchmarks at a fraction of the cost, establishing it as an essential baseline for evaluating novel fine-tuning techniques. Third, our analysis reveals that substantial priors can be learned solely from the text modality and that using a textual caption instead of the video yields performance on par with Vanilla SFT. These surprising findings reveal the limitations of current MLLMs when it comes to social understanding. IntentBench-Prime, Vanilla SFT model, and code are publicly available.
Problem

Research questions and friction points this paper is trying to address.

social understanding
audio-visual question answering
multimodal large language models
reasoning
benchmark noise
Innovation

Methods, ideas, or system contributions that make the work stand out.

IntentBench-Prime
Vanilla SFT
multimodal reasoning
social understanding
audio-visual QA
💼 Related Jobs
No related jobs found.