Beyond Worst-Case Coreset Bounds for $k$-Clustering via Determinantal Sampling

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对大规模数据集的聚类任务,通过引入基于行列式点过程的新型相关采样框架,解决了构建更小ε-coreset的问题,超越了最坏情况下的理论界限。
📝 Abstract
Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clustering tasks where memory and computational constraints demand compact yet faithful summaries. A standard approach is to construct an \textit{$\epsilon$-coreset}: a small weighted subset that approximately preserves the clustering cost for every plausible choice of centers. For the \textit{$(k,z)$-clustering problem}, existing worst-case bounds on coreset size are essentially tight, ruling out substantially smaller coresets in general. However, such worst-case instances are often unrepresentative of real-world data. In this work, we show that significantly smaller coresets are possible under mild and natural assumptions on the underlying data distribution. We introduce a new correlated sampling framework, called \textit{determinantal sampling}, based on a novel application of determinantal point processes. Using this framework, we obtain an efficiently constructible $\varepsilon$-coreset for $(k,z)$-clustering in $\mathbb R^d$ whose dependence on $1/\varepsilon$ has exponent strictly smaller than $2$ when $d$ is fixed. This improves over the worst-case $\varepsilon^{-2}$ barrier under our beyond-worst-case assumptions. To the best of our knowledge, this is the first result that provably surpasses these lower bounds through beyond-worst-case assumptions. Finally, we validate our approach on synthetic and real-world benchmark datasets, where it consistently achieves smaller coresets than existing state-of-the-art methods, even without explicitly enforcing the assumptions used in the analysis.
Problem

Research questions and friction points this paper is trying to address.

data reduction
clustering
ε-coreset
worst-case bounds
determinantal sampling
Innovation

Methods, ideas, or system contributions that make the work stand out.

determinantal sampling
beyond-worst-case assumptions
ε-coreset
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Diptarka Chakraborty
Diptarka Chakraborty
School of Computing, National University of Singapore
Theoretical Computer Science
Satyaki Mukherjee
Satyaki Mukherjee
National University of Singapore
ProbabilityRandom Matrix TheoryMachine learning
G
Gaurav Vallabhdas Revankar
Department of Mathematics, Indian Institute of Technology Bombay
H
Hoang-Son Tran
Department of Mathematics, National University of Singapore