Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing audio language models in modeling spatial location and motion trajectories of sound sources, as well as the absence of spatial audio cues in vision-language models, which hinders joint identification, localization, and tracking of dynamic sound sources. To bridge this gap, the authors introduce ST-OmniQA, the first spatiotemporal audio-visual question answering benchmark that integrates 360° video with first-order Ambisonics audio, along with a multimodal model, ST-Omni-R1, which fuses semantic, trajectory, and visual contextual information for reasoning. They further propose a novel four-level evaluation protocol encompassing sound source identification, azimuth, distance, and motion trajectory, combined with progressive curriculum learning and reasoning-tree-based reinforcement learning to achieve cross-modal spatiotemporal alignment. Experiments show that the proposed approach achieves a semantic accuracy of 77.83% on ST-OmniQA—substantially outperforming baselines (37.28%)—and demonstrates strong transferability across multiple established spatial audio benchmarks.
📝 Abstract
Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources. To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering benchmark built from panoramic videos paired with synchronized first-order Ambisonics (FOA) audio of moving sound sources. It contains 40K videos and 400K question-answer pairs organized into four capability levels covering sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded audio-visual reasoning. Building on this benchmark, we propose ST-Omni-R1, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning. ST-Omni-R1 achieves 77.83\% average semantic accuracy across the four levels, compared with 37.28\% for the best evaluated baseline. Results on three public spatial-audio benchmarks further indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.
Problem

Research questions and friction points this paper is trying to address.

audio-visual reasoning
sound event localization
spatio-temporal tracking
omni-modal language models
spatial audio
Innovation

Methods, ideas, or system contributions that make the work stand out.

spatio-temporal audio-visual reasoning
first-order Ambisonics
omni-modal language models
sound source localization and tracking
curriculum learning
🔎 Similar Papers
No similar papers found.