摘要
With a large number of distance measures, the appropriate choice for clustering a given data set with a specified clustering algorithm becomes an important problem. In this article, an automatic distance measure recommendation method for clustering algorithms is proposed. The recommendation method consists of the following steps: (1) metadata extraction, including meta-feature collection and meta-target identification; (2) recommendation model construction using metadata; and (3) distance measure recommendation for a new data set by the recommendation model. Two different types of meta-targets and meta-learning techniques are utilized considering the possible different requirements of users. To validate the necessity and effectiveness of the distance measure recommendation method, an empirical study is conducted with 199 publicly available data sets, 9 distance measures, and 2 widely used clustering algorithms. The experimental results indicate that distance measure significantly influences the performance of the clustering algorithm for a given data set. Furthermore,performance analysis of the proposed recommendation method proves its effectiveness.
| 源语言 | 英语 |
|---|---|
| 期刊论文编号 | 7 |
| 期刊 | ACM Transactions on Knowledge Discovery from Data |
| 卷 | 15 |
| 期 | 1 |
| DOI | |
| 出版状态 | 已出版 - 1月 2021 |
学术指纹
探究 'Automatic Recommendation of a Distance Measure for Clustering Algorithms' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver