Estimating the Optimal Number of Clusters in Categorical Data Clustering by Silhouette Coefficient
Addressing the challenge of automatic optimal cluster number selection in categorical data clustering, this paper introduces the first unsupervised, parameter-free method that adapts the silhouette coefficient to purely categorical spaces. The core methodological contribution lies in designing a dissimilarity measure grounded in Hamming distance and attribute matching, reformulating intra- and inter-cluster silhouette computations accordingly, and establishing a comprehensive clustering evaluation framework tailored specifically for categorical data. Extensive experiments across multiple real-world categorical datasets demonstrate that the proposed approach significantly outperforms mainstream indices—including the Calinski–Harabasz (CH) index and Gap Statistic—reducing average cluster number estimation error by 37%. The method exhibits strong robustness, intrinsic interpretability, and broad generalizability across diverse categorical domains.