Abstract
A method of discovering new Chinese words from Internet based on information propagation is proposed to solve the problems that the recognizing results of existing methods always have short life cycles and will not be used again in soon. The method combines the characteristics of new words such as widely spreading and long lasting, and three statistics, i. e. coverage rate of users, coverage rate of topics and life cycle of a new word, are defined. The N-gram algorithm is applied to generate candidates of new words, then the word candidates are filtered bade on word frequency and word flexibility. Experiments with the text of microblogs as corpus and comparisons with the existing methods show that the user statistic enhances the accuracy rate of recognizing new words by 11%, the topic statistic enhances the accuracy rate by 10%, and the time statistic enhances the accuracy rate by 13%. When the three statistics are combined, the accuracy rate is raised by 16%. It can be concluded that each single statistic considered by the proposed method can enhance the accuracy rate, and more accurate rate can be obtained by considering the combination of the three statistics rather than just considering one statistic.
| Original language | English |
|---|---|
| Pages (from-to) | 59-64 |
| Number of pages | 6 |
| Journal | Hsi-An Chiao Tung Ta Hsueh/Journal of Xi'an Jiaotong University |
| Volume | 49 |
| Issue number | 12 |
| DOIs | |
| State | Published - 10 Dec 2015 |
Keywords
- Information propagation
- New word discovery
- Temporal characteristics
- User behavior
Fingerprint
Dive into the research topics of 'A method of discovering new chinese words from internet based on information propagation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver