摘要
Focused crawlers aim to search only a subset of the web related to a specific topic whose performance mainly depends on the accuracy of predicting the relevance of a newly seen URL. In this paper a novel focused crawling method based on sequential probabilistic model and reinforcement learning is proposed. The crawled website is modeled as a graph and a model of Conditional Random Fields is trained over it which learns the optimal paths that lead to relevant pages on the websites through exploit a variety of features in the hyperlinks, HTML tags, and page sequence. Then, a reinforcement learning algorithm is used to map each hyperlink on the crawl frontier to a future discounted reward as its priority. Experimental results using a large number of web pages from diverse domains show that our technique provides better performance than traditional focused crawlers.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 1657-1664 |
| 页数 | 8 |
| 期刊 | Journal of Computational Information Systems |
| 卷 | 3 |
| 期 | 4 |
| 出版状态 | 已出版 - 4月 2007 |
学术指纹
探究 'Probabilistic graphical model for efficient focused web crawling' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver