Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过增加帧级定位模型来解决大型音频-语言模型在精细时间感知上的局限,提高了事件定位的精度和可靠性。
📝 Abstract
Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal perception, particularly in precise event localization. Existing approaches primarily post-train LALMs to predict event boundaries as timestamp tokens. However, this generative formulation lacks explicit correspondence between the timestamp predictions and fine-grained acoustic evidence, limiting the precision and reliability of temporal localization. To address this issue, we augment the LALM with a dedicated frame-level grounding model while leveraging its semantic modeling capability to represent the event query. Specifically, the frozen LALM encodes the event query with audio as context, and the grounding model combines these query representations with fine-grained audio features to localize the target event at the frame level. Extensive experiments across diverse temporal grounding benchmarks demonstrate strong and consistent improvements over existing methods. Further evaluation shows that the grounding model can provide temporal evidence to support downstream reasoning.
Problem

Research questions and friction points this paper is trying to address.

Large Audio-Language Models
fine-grained temporal perception
event localization
timestamp predictions
acoustic evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

frame-level grounding
fine-grained temporal perception
event localization
💼 Related Jobs
No related jobs found.
Y
Yanfeng Shi
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
Y
Yan Song
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
J
Junhui Li
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
T
Tinggan Huang
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
W
Wu Guo
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
H
Haoyu Song
ICT Cluster, Singapore Institute of Technology, Singapore
Ian McLoughlin
Ian McLoughlin
Professor Singapore Institute of Technology (Singapore) and USTC (China)
AI for speech & audiosignal processingembedded systemscomputer architecture