🤖 AI Summary
Addressing the challenge of automatic optimal cluster number selection in categorical data clustering, this paper introduces the first unsupervised, parameter-free method that adapts the silhouette coefficient to purely categorical spaces. The core methodological contribution lies in designing a dissimilarity measure grounded in Hamming distance and attribute matching, reformulating intra- and inter-cluster silhouette computations accordingly, and establishing a comprehensive clustering evaluation framework tailored specifically for categorical data. Extensive experiments across multiple real-world categorical datasets demonstrate that the proposed approach significantly outperforms mainstream indices—including the Calinski–Harabasz (CH) index and Gap Statistic—reducing average cluster number estimation error by 37%. The method exhibits strong robustness, intrinsic interpretability, and broad generalizability across diverse categorical domains.