TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
This work addresses the scarcity of large-scale multilingual read-speech datasets with reliable speaker labels, which has hindered progress in tasks such as anti-spoofing and speaker verification. We present the first systematic resolution of speaker label heterogeneity in Mozilla Common Voice, introducing TidySpeech—a high-quality dataset comprising over 210,000 monolingual and 4,500 multilingual speakers. We further define standardized evaluation protocols for monolingual (Tidy-M) and multilingual (Tidy-X) scenarios. A ResNet-based model fine-tuned on Tidy-M achieves an equal error rate (EER) of 0.35% and demonstrates significantly improved generalization on the unseen conversational dataset CANDOR. Both the dataset and models are publicly released, establishing the first large-scale open benchmark for cross-lingual speaker verification.