The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of multi-speaker speech recognition and understanding in egocentric smart glasses scenarios, where dynamic acoustic conditions, speaker overlap, and spatial ambiguity pose significant difficulties. To this end, the authors introduce the first large-scale, real-world evaluation benchmark tailored for smart glasses, comprising 106 hours of four-channel near-ear microphone array recordings, with dedicated tracks for dyadic conversations and multi-party meetings. The benchmark jointly evaluates timestamped speaker-attributed automatic speech recognition (TSA-ASR) and spoken language understanding (SLU). Experimental results reveal that severe speaker overlap substantially degrades TSA-ASR accuracy, and current audio-language models still exhibit limited capacity in leveraging paralinguistic cues for complex SLU tasks, underscoring the benchmark’s value in advancing research on wearable multimodal interaction.
📝 Abstract
Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have created new opportunities for wearable speech interfaces, with smart glasses providing an egocentric platform for continuous audio sensing and assistance. However, speech recognition and understanding in this setting remain challenging because of dynamic acoustic conditions, speaker overlap, and the spatial ambiguity introduced by wearer-centered recording geometry. To support systematic evaluation in this setting, we introduce the IEEE SLT 2026 SmartGlasses Challenge for egocentric multi-speaker speech processing. The challenge consists of two tracks, Dyadic Dialogue Understanding and Multi-party Meeting Understanding, and jointly evaluates Time-Stamped Speaker-Attributed Automatic Speech Recognition (TSA-ASR) and Spoken Language Understanding (SLU). It is built on a 106-hour four-channel egocentric speech dataset containing 714 sessions collected in real-world scenarios. This paper describes challenge tasks, dataset construction, submissions, and summarizes the main findings from the shared evaluation. The results show that heavy speaker overlap remains a major factor affecting TSA-ASR performance, while paralinguistic acoustic understanding continues to be difficult for current audio-language models in complex SLU settings. Further details can be found on the official challenge website.
Problem

Research questions and friction points this paper is trying to address.

egocentric speech
multi-talker speech recognition
speaker overlap
spatial ambiguity
spoken language understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

egocentric speech
TSA-ASR
audio-language models
multi-talker speech recognition
spoken language understanding