Skip to main navigation Skip to search Skip to main content

Self-Enhanced Density Clustering for High Dimension and Low Sample Size Data

  • Bingbing Jiang
  • , Zhongli Wang
  • , Jie Yang
  • , Guang Kui Xu
  • , Wei Chen
  • , Chenglong Zhang
  • , Xinyan Liang
  • , Peng Zhou
  • , Weiguo Sheng
  • , Weiping Ding
  • Hangzhou Normal University
  • The University of Sydney
  • Nanjing University
  • Shanxi University
  • Anhui University
  • Nantong University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

Clustering on high-dimensional and low sample size (HDLSS) data remains a critical, persistent challenge where extreme sparsity and noise confound cluster analysis. This creates a dilemma: spectral methods fail as distance metrics degrade, while deep clustering tends to over-fit scarce data. To break this dilemma, a Self-Enhanced Density Clustering (SEDC) framework that integrates the cluster structure discovery and embedding representation learning into an iterative enhancement process is proposed in this paper. Specifically, SEDC uses adaptive density-derived centroids to parameterize probabilistic soft labels, which in turn supervise a lightweight multilayer perceptron (MLP) to learn the low-dimensional embedding from data. The resulting embedding provides a refined metric space for further generating superior labels in the subsequent interaction process. This feedback forms a mutual reinforcement that progressively enhances the discrimination of embedding while rigorously mitigating over-fitting. Extensive experiments on 43 challenging HDLSS datasets demonstrate state-of-the-art performance, substantially outperforming popular clustering methods. This work delivers a principled and promising solution for robust data clustering in HDLSS situations.

Original languageEnglish
Title of host publicationKDD 2026 - Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1
PublisherAssociation for Computing Machinery
Pages508-519
Number of pages12
ISBN (Electronic)9798400722585
DOIs
StatePublished - 20 Apr 2026
Event32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026 - Jeju Island, Korea, Republic of
Duration: 9 Aug 202613 Aug 2026

Publication series

NameProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
Volume1-A
ISSN (Print)2154-817X

Conference

Conference32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026
Country/TerritoryKorea, Republic of
CityJeju Island
Period9/08/2613/08/26

Keywords

  • clustering
  • density estimation
  • gene expression data analysis
  • high-dimensional data
  • low sample size data

Fingerprint

Dive into the research topics of 'Self-Enhanced Density Clustering for High Dimension and Low Sample Size Data'. Together they form a unique fingerprint.

Cite this