Pushing the Limits of End-to-End Diarization
This study addresses the limited accuracy of joint segmentation and clustering in end-to-end speaker diarization. We propose a unified solution based on an extended non-autoregressive EEND-TA architecture. To enhance modeling capability for complex overlapping speech and long-tail scenarios, we construct a large-scale synthetic dataset covering up to eight simultaneous speakers and adopt a pretraining-finetuning paradigm to improve generalization. Our method directly outputs speaker label sequences without requiring post-processing steps such as clustering or voice activity detection (VAD), thereby simplifying the pipeline and improving robustness. Evaluated on standard benchmarks—including AliMeeting, AMI, and DIHARD III—our approach achieves state-of-the-art performance, with a DER of 14.49% on DIHARD III, demonstrating strong cross-domain adaptability and practical deployability.