摘要
Clustering is one of fundamental tasks of data mining and machine learning. Due to the limitation of cluster assumption, lots of clustering algorithms perform poorly on some datasets against their assumptions, especially high-dimensional big data. This paper presents a maximum average entropy-rate based correlation clustering algorithm which is a kind of a graph-based correlation clustering. The objective function of original correlation clustering is decomposed into several single cluster optimizations and the limitation of big data in correlation clustering is removed by the neighboring connected graph. In algorithm implementation, the optimization of proper neighbor searching and correlation clustering are performed by heuristic neighbor searching and cluster generating respectively, and there is also an efficient graph-iterated implementation on distributed computation platform. Compared with other clustering algorithms, the proposed clustering algorithm is more flexible in cluster assumption, when accelerating the clustering process. In an experimental study we demonstrate the performance of the proposed algorithms on several datasets. The proposed clustering algorithm performed better than the other six clustering algorithms on the highest f1-measure and purity values, while its running time on high-dimensional big data is much lower than other clustering algorithms.
| 投稿的翻译标题 | Maximum average entropy-rate based correlation clustering for big data |
|---|---|
| 源语言 | 繁体中文 |
| 页(从-至) | 1572-1585 |
| 页数 | 14 |
| 期刊 | Scientia Sinica Informationis |
| 卷 | 49 |
| 期 | 12 |
| DOI | |
| 出版状态 | 已出版 - 2019 |
关键词
- big data
- clustering
- correlation clustering
- entropy-rate
- graph-based clustering
学术指纹
探究 '基于最大平均熵率的大数据关联聚类算法' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver