Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets
This work proposes an efficient method for evaluating the intrinsic quality (IQ) of large-scale face recognition datasets without requiring full model training. By integrating neighborhood consistency scores with the effective rank of the embedding space, the approach establishes a lightweight, validation-free quality assessment framework capable of rapidly predicting downstream recognition performance using proxy models or dataset subsets. Experimental results demonstrate that the proposed IQ metric accurately forecasts model performance across clean, noisy, and mixed-quality datasets, substantially reducing the cost of data diagnosis and filtering. This provides a practical and scalable tool for preprocessing massive face datasets in real-world applications.