Persistent Multiscale Density-based Clustering

📅 2025-12-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Density-based clustering often suffers from sensitivity to manually tuned hyperparameters—such as density thresholds or minimum cluster size—especially when prior knowledge about data distribution is unavailable, leading to poor robustness. To address this, we propose PLSCAN, the first algorithm integrating scale-space clustering with persistent homology theory to construct an adaptive metric space that automatically identifies stable leaf clusters of HDBSCAN* across all scales—without any hyperparameter tuning. Our method leverages hierarchical density estimation via mutual reachability distance, scale-space trajectory tracking, and extraction of persistent leaf nodes. Experiments demonstrate that PLSCAN achieves higher average Adjusted Rand Index (ARI) than HDBSCAN* on multiple real-world datasets and exhibits superior robustness to perturbations in neighborhood size. Its computational efficiency is comparable to k-Means in low dimensions and matches HDBSCAN* in high dimensions. The core contribution is a fully parameter-free framework for extracting multi-scale density structures.

Technology Category

Application Category

📝 Abstract
Clustering is a cornerstone of modern data analysis. Detecting clusters in exploratory data analyses (EDA) requires algorithms that make few assumptions about the data. Density-based clustering algorithms are particularly well-suited for EDA because they describe high-density regions, assuming only that a density exists. Applying density-based clustering algorithms in practice, however, requires selecting appropriate hyperparameters, which is difficult without prior knowledge of the data distribution. For example, DBSCAN requires selecting a density threshold, and HDBSCAN* relies on a minimum cluster size parameter. In this work, we propose Persistent Leaves Spatial Clustering for Applications with Noise (PLSCAN). This novel density-based clustering algorithm efficiently identifies all minimum cluster sizes for which HDBSCAN* produces stable (leaf) clusters. PLSCAN applies scale-space clustering principles and is equivalent to persistent homology on a novel metric space. We compare its performance to HDBSCAN* on several real-world datasets, demonstrating that it achieves a higher average ARI and is less sensitive to changes in the number of mutual reachability neighbours. Additionally, we compare PLSCAN's computational costs to k-Means, demonstrating competitive run-times on low-dimensional datasets. At higher dimensions, run times scale more similarly to HDBSCAN*.
Problem

Research questions and friction points this paper is trying to address.

Automates hyperparameter selection for density-based clustering
Identifies stable clusters across varying minimum size thresholds
Improves robustness and performance over existing clustering methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

PLSCAN algorithm identifies stable HDBSCAN* leaf clusters
Applies scale-space clustering and persistent homology principles
Achieves higher ARI and lower sensitivity to parameters
🔎 Similar Papers
No similar papers found.
D
Daniël Bot
UHasselt, Data Science Institute (DSI)
L
Leland McInnes
Tutte Institute for Mathematics and Computing
Jan Aerts
Jan Aerts
KU Leuven
visual analyticsdata visualizationbioinformaticsgenomicsagriculture