🤖 AI Summary
K-means clustering is frequently employed in psychology to identify latent subgroups; however, its reliance on geometric distance precludes validation of whether the resulting clusters correspond to genuine psychological constructs. This study systematically compares K-means performance on multidimensional Gaussian simulated data—lacking any true categorical structure—with that on the international psychometric dataset SMARVUS. The results demonstrate that K-means consistently produces stable and visually coherent clusters even in continuous latent variable spaces devoid of discrete classes. These findings suggest that such clusters may merely reflect spatial partitioning rather than meaningful psychological types, thereby challenging the foundational assumption that K-means can validly infer the existence of latent categories in psychological research.
📝 Abstract
K-means clustering is widely used in psychological and psychometric research to identify profiles, subgroups, and potential typologies, yet its classical formulation does not test whether such groups exist as latent psychological categories. Instead, K-means partitions multidimensional space into regions around centroids, favoring compact, approximately spherical clusters defined by geometric distance. In this paper, we examine this limitation through a sequence of controlled simulated datasets. We then extend the analysis to the SMARVUS dataset, a large international psychometric dataset comprising survey responses from university students across 35 countries, to evaluate whether similar geometric partitioning patterns emerge in empirical psychological data. By contrasting simulated and empirical data, this paper argues that K-means can produce stable and visually coherent clustering solutions even in continuous Gaussian latent spaces without true subgroup structure.