Poisson Subspace Clustering: Focusing on the Essentials in Count Data

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对计数数据聚类问题,提出基于泊松分布的3CPO算法,通过最大化后验概率来找到高质量的聚类解,并增强结果解释性。
📝 Abstract
Count data represented as a matrix of non-negative integer values, such as contingency tables, are prevalent across diverse domains. When clustering such data sets, specific methods are required, as generic algorithms often fail to consider their unique distributional properties, leading to unreliable outputs. An effective strategy is to use well-established statistical models such as the Poisson and negative binomial distributions. We present 3CPO, a clustering algorithm based on statistically solid modeling of count data. In addition to the cluster labels, it identifies a subset of relevant columns, enhancing the interpretability of the results. We propose a simple iterative algorithm that maximizes the posterior probability to find good clustering solutions and discuss its properties. Extensive experiments demonstrate its ability to define high-quality clusters within associated subspaces for various data domains, ranging from gene expressions and texts to economics. Our findings suggest that 3CPO is a robust solution for clustering count data in a statistically sound and interpretable manner. Our code is available at https://github.com/collinleiber/3CPO.
Problem

Research questions and friction points this paper is trying to address.

count data
clustering
Poisson distribution
statistical models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Poisson Subspace Clustering
count data
statistical models
interpretability
posterior probability
🔎 Similar Papers
2024-09-01arXiv.orgCitations: 4