Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决长音频会议理解中声学信息丢失和长期上下文记忆差的问题,构建了LongAudioQA数据集,并提出GRGA模型,将异构音频特征建模成多维图并通过代理规划进行检索和答案生成。
📝 Abstract
While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Existing speech QA paradigms and state-of-the-art Speech LLMs suffer from acoustic information loss and poor long-term context memory. To address these issues, we construct the LongAudioQA dataset and propose the GRGA model, which models heterogeneous audio features into a multi-dimensional graph and leverages agent planning for retrieval and answer generation.
Problem

Research questions and friction points this paper is trying to address.

Long-form audio meeting understanding
Question answering
Acoustic information loss
Long-term context memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

LongAudioQA
GRGA model
heterogeneous audio features
agent planning