🤖 AI Summary
This work addresses the challenge that single-bandwidth kernel spectral clustering struggles to capture the complex multiscale distance structures inherent in high-dimensional data. To overcome this limitation, the authors propose a multi-kernel spectral clustering method that adaptively selects bandwidths via empirical quantiles and fuses multiple kernel functions to construct a low-rank approximate similarity matrix. They establish a row-wise ℓ₂,∞ perturbation bound for the leading eigenvectors of the normalized Laplacian, enabling fine-grained control over observation-level spectral embeddings. Under conditions of sufficient eigenvalue gap and inter-cluster separation, the proposed approach—combined with approximate K-means—achieves exact cluster recovery with high probability, significantly improving clustering performance in high-dimensional mixture models with multiscale structure.
📝 Abstract
Kernel spectral clustering with a single bandwidth can be inadequate for data exhibiting multiple characteristic pairwise-distance scales, a problem particularly prevalent in the high-dimensional regime. We address this issue through a multi-kernel formulation that aggregates kernels with different bandwidths. The bandwidths are selected as prescribed empirical quantiles of the pairwise squared distances, thereby capturing the relevant distance scales without requiring prior population-scale information.
We develop a rigorous theoretical analysis of the resulting method under a general high-dimensional, multi-scale mixture model with heterogeneous cluster centers and covariance geometries. We construct a blockwise constant, low-rank informative approximation to the empirical multi-kernel matrix and establish row-wise $\ell_{2,\infty}$ perturbation bounds for its leading spectral components, as well as for the associated normalized Laplacian matrix. These bounds yield observation-level control of the spectral embedding, which is more informative than conventional global eigenspace perturbation estimates. Under suitable eigen-gap and cluster-separation conditions, we show that approximate $K$-means applied to the multi-kernel spectral embedding achieves exact recovery with high probability.