TY - GEN
T1 - Self-Enhanced Density Clustering for High Dimension and Low Sample Size Data
AU - Jiang, Bingbing
AU - Wang, Zhongli
AU - Yang, Jie
AU - Xu, Guang Kui
AU - Chen, Wei
AU - Zhang, Chenglong
AU - Liang, Xinyan
AU - Zhou, Peng
AU - Sheng, Weiguo
AU - Ding, Weiping
N1 - Publisher Copyright:
© 2026 Owner/Author.
PY - 2026/4/20
Y1 - 2026/4/20
N2 - Clustering on high-dimensional and low sample size (HDLSS) data remains a critical, persistent challenge where extreme sparsity and noise confound cluster analysis. This creates a dilemma: spectral methods fail as distance metrics degrade, while deep clustering tends to over-fit scarce data. To break this dilemma, a Self-Enhanced Density Clustering (SEDC) framework that integrates the cluster structure discovery and embedding representation learning into an iterative enhancement process is proposed in this paper. Specifically, SEDC uses adaptive density-derived centroids to parameterize probabilistic soft labels, which in turn supervise a lightweight multilayer perceptron (MLP) to learn the low-dimensional embedding from data. The resulting embedding provides a refined metric space for further generating superior labels in the subsequent interaction process. This feedback forms a mutual reinforcement that progressively enhances the discrimination of embedding while rigorously mitigating over-fitting. Extensive experiments on 43 challenging HDLSS datasets demonstrate state-of-the-art performance, substantially outperforming popular clustering methods. This work delivers a principled and promising solution for robust data clustering in HDLSS situations.
AB - Clustering on high-dimensional and low sample size (HDLSS) data remains a critical, persistent challenge where extreme sparsity and noise confound cluster analysis. This creates a dilemma: spectral methods fail as distance metrics degrade, while deep clustering tends to over-fit scarce data. To break this dilemma, a Self-Enhanced Density Clustering (SEDC) framework that integrates the cluster structure discovery and embedding representation learning into an iterative enhancement process is proposed in this paper. Specifically, SEDC uses adaptive density-derived centroids to parameterize probabilistic soft labels, which in turn supervise a lightweight multilayer perceptron (MLP) to learn the low-dimensional embedding from data. The resulting embedding provides a refined metric space for further generating superior labels in the subsequent interaction process. This feedback forms a mutual reinforcement that progressively enhances the discrimination of embedding while rigorously mitigating over-fitting. Extensive experiments on 43 challenging HDLSS datasets demonstrate state-of-the-art performance, substantially outperforming popular clustering methods. This work delivers a principled and promising solution for robust data clustering in HDLSS situations.
KW - clustering
KW - density estimation
KW - gene expression data analysis
KW - high-dimensional data
KW - low sample size data
UR - https://www.scopus.com/pages/publications/105038096176
U2 - 10.1145/3770854.3780293
DO - 10.1145/3770854.3780293
M3 - 会议稿件
AN - SCOPUS:105038096176
T3 - Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
SP - 508
EP - 519
BT - KDD 2026 - Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1
PB - Association for Computing Machinery
T2 - 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026
Y2 - 9 August 2026 through 13 August 2026
ER -