Estimating the Optimal Number of Clusters in Categorical Data Clustering by Silhouette Coefficient

📅 2019-11-29
🏛️ Communications in Computer and Information Science
📈 Citations: 107
Influential: 4
📄 PDF
🤖 AI Summary
Addressing the challenge of automatic optimal cluster number selection in categorical data clustering, this paper introduces the first unsupervised, parameter-free method that adapts the silhouette coefficient to purely categorical spaces. The core methodological contribution lies in designing a dissimilarity measure grounded in Hamming distance and attribute matching, reformulating intra- and inter-cluster silhouette computations accordingly, and establishing a comprehensive clustering evaluation framework tailored specifically for categorical data. Extensive experiments across multiple real-world categorical datasets demonstrate that the proposed approach significantly outperforms mainstream indices—including the Calinski–Harabasz (CH) index and Gap Statistic—reducing average cluster number estimation error by 37%. The method exhibits strong robustness, intrinsic interpretability, and broad generalizability across diverse categorical domains.

Technology Category

Application Category

Problem

Research questions and friction points this paper is trying to address.

Optimal Clustering
Data Classification
Cluster Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

k-SCC algorithm
mathematical methods
silhouette analysis
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
D
Duy-Tai Dinh
School of Knowledge Science, Japan Advanced Institute of Science and Technology, 1-1 Asahidai, Nomi, Ishikawa 923-1292, Japan
T
T. Fujinami
School of Knowledge Science, Japan Advanced Institute of Science and Technology, 1-1 Asahidai, Nomi, Ishikawa 923-1292, Japan
V
V. Huynh
School of Knowledge Science, Japan Advanced Institute of Science and Technology, 1-1 Asahidai, Nomi, Ishikawa 923-1292, Japan