AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出AudioLens-R1模型,通过推理蒸馏和偏好优化训练,解决多视角语音聚类问题,提高音频集合的组织灵活性。
📝 Abstract
Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.
Problem

Research questions and friction points this paper is trying to address.

audio clustering
multi-perspective
speech collections
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-perspective clustering
audio-language model
reasoning distillation
preference optimization
AudioLens-Bench
🔎 Similar Papers
No similar papers found.