Skip to main navigation Skip to search Skip to main content

基于最大平均熵率的大数据关联聚类算法

Translated title of the contribution: Maximum average entropy-rate based correlation clustering for big data
  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

Abstract

Clustering is one of fundamental tasks of data mining and machine learning. Due to the limitation of cluster assumption, lots of clustering algorithms perform poorly on some datasets against their assumptions, especially high-dimensional big data. This paper presents a maximum average entropy-rate based correlation clustering algorithm which is a kind of a graph-based correlation clustering. The objective function of original correlation clustering is decomposed into several single cluster optimizations and the limitation of big data in correlation clustering is removed by the neighboring connected graph. In algorithm implementation, the optimization of proper neighbor searching and correlation clustering are performed by heuristic neighbor searching and cluster generating respectively, and there is also an efficient graph-iterated implementation on distributed computation platform. Compared with other clustering algorithms, the proposed clustering algorithm is more flexible in cluster assumption, when accelerating the clustering process. In an experimental study we demonstrate the performance of the proposed algorithms on several datasets. The proposed clustering algorithm performed better than the other six clustering algorithms on the highest f1-measure and purity values, while its running time on high-dimensional big data is much lower than other clustering algorithms.

Translated title of the contributionMaximum average entropy-rate based correlation clustering for big data
Original languageChinese (Traditional)
Pages (from-to)1572-1585
Number of pages14
JournalScientia Sinica Informationis
Volume49
Issue number12
DOIs
StatePublished - 2019

Fingerprint

Dive into the research topics of 'Maximum average entropy-rate based correlation clustering for big data'. Together they form a unique fingerprint.

Cite this