Cluster-Based Dimensionality Reduction by Nonparametric Distributional Screening

๐Ÿ“… 2026-09-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
่ฏฅ็ ”็ฉถ้€š่ฟ‡้žๅ‚ๆ•ฐๅˆ†ๅธƒ็ญ›้€‰ๆ–นๆณ•๏ผŒ้’ˆๅฏน้ซ˜็ปดๆ•ฐๆฎๅŠๅ…ถ้ข„ๅฎšไน‰็š„่š็ฑป๏ผŒ้€‰ๆ‹ฉไฟ็•™ๅŽŸๅง‹ๅๆ ‡ๅญ้›†ไปฅไฟๆŒๅŒบๅˆ†่š็ฑป็š„ๅˆ†ๅธƒไฟกๆฏใ€‚
๐Ÿ“ Abstract
We consider dimensionality reduction for high-dimensional observations accompanied by a supplied partition into two or more clusters. The objective is not to construct a low-rank projection, but to retain an interpretable subset of the original coordinates that preserves the distributional information distinguishing the clusters. For each coordinate, the proposed procedure compares the cluster-specific empirical distribution functions through a several-sample Kolmogorov-Smirnov separation statistic. We formalize the resulting marginal cluster support and establish simultaneous finite-sample concentration over all coordinates, explicit bounds for false inclusions and omissions, and exact support recovery when the minimum distributional separation dominates the high-dimensional stochastic error. We also quantify the dimension inflation induced by using an unadjusted testing level and give a familywise-error-controlled version. Under a conditional sufficiency condition, sure screening preserves the full-data posterior cluster probabilities, mutual information, and Bayes risk; an additional result characterizes robustness to imperfectly estimated cluster labels. The procedure is invariant to strictly increasing coordinate transformations and can retain low-variance cluster signals that principal components may discard. We further develop average dual information, a criterion combining partition agreement after transformation with structural coverage of cluster-relevant coordinates, and derive its basic properties and consistency. Simulations illustrate the theory, the interpretability of the selected coordinates, and the distinction between cluster-directed screening and variance-directed projection.
Problem

Research questions and friction points this paper is trying to address.

Dimensionality Reduction
Cluster Analysis
Distributional Information
Kolmogorov-Smirnov
High-dimensional Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nonparametric Distributional Screening
Cluster-Specific Empirical Distributions
Kolmogorov-Smirnov Separation Statistic
Marginal Cluster Support
Dimensionality Reduction
๐Ÿ”Ž Similar Papers
2021-06-14IEEE Transactions on Visualization and Computer GraphicsCitations: 12
๐Ÿ’ผ Related Jobs
No related jobs found.