Drawing Lines in Psychological Space: What K-means Clustering Reveals in Simulated and Real Psychometric Data

📅 2026-05-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
K-means clustering is frequently employed in psychology to identify latent subgroups; however, its reliance on geometric distance precludes validation of whether the resulting clusters correspond to genuine psychological constructs. This study systematically compares K-means performance on multidimensional Gaussian simulated data—lacking any true categorical structure—with that on the international psychometric dataset SMARVUS. The results demonstrate that K-means consistently produces stable and visually coherent clusters even in continuous latent variable spaces devoid of discrete classes. These findings suggest that such clusters may merely reflect spatial partitioning rather than meaningful psychological types, thereby challenging the foundational assumption that K-means can validly infer the existence of latent categories in psychological research.
📝 Abstract
K-means clustering is widely used in psychological and psychometric research to identify profiles, subgroups, and potential typologies, yet its classical formulation does not test whether such groups exist as latent psychological categories. Instead, K-means partitions multidimensional space into regions around centroids, favoring compact, approximately spherical clusters defined by geometric distance. In this paper, we examine this limitation through a sequence of controlled simulated datasets. We then extend the analysis to the SMARVUS dataset, a large international psychometric dataset comprising survey responses from university students across 35 countries, to evaluate whether similar geometric partitioning patterns emerge in empirical psychological data. By contrasting simulated and empirical data, this paper argues that K-means can produce stable and visually coherent clustering solutions even in continuous Gaussian latent spaces without true subgroup structure.
Problem

Research questions and friction points this paper is trying to address.

K-means clustering
psychological categories
latent structure
psychometric data
subgroup detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

K-means clustering
latent psychological categories
simulated data
psychometric data
cluster validity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pedro Henrique Ramos Pinto
Postgraduate Program in Cognitive and Behavioral Neuroscience (PPGNeC) - Federal University of Paraíba (UFPB) - João Pessoa, PB - Brazil
M
Maria Jullyanna Ferreira Marques
Center for Health Sciences (CCS) - Federal University of Paraíba (UFPB) - João Pessoa, PB - Brazil
L
Luiz Carlos Serramo Lopez
Postgraduate Program in Cognitive and Behavioral Neuroscience (PPGNeC) - Federal University of Paraíba (UFPB) - João Pessoa, PB - Brazil