跳到主要导航 跳到搜索 跳到主要内容

Learning Deep Web crawling with diverse features

  • Xi'an Jiaotong University

科研成果: 书/报告/会议事项章节会议稿件同行评审

18 引用 (Scopus)

摘要

The key to Deep Web crawling is to submit promising keywords to query form and retrieve Deep Web content efficiently. To select keywords, existing methods make a decision based on keywords' statistic information deriving from TF and DF in local acquired records, thus work well only in textual databases providing full text search interfaces, whereas not well in structured databases of multi-attribute or field-restricted search interfaces. This paper proposes a novel Deep Web crawling method. Keywords are encoded as a tuple by its linguistic, statistic and HTML features so that a harvest rate evaluation model can be learned from the issued keywords for the un-issued in future. The method breaks through the assumption of plain-text search made by existing methods. Experimental results show that the method outperforms the state of the art methods.

源语言英语
主期刊名Proceedings - 2009 IEEE/WIC/ACM International Conference on Web Intelligence, WI 2009
572-575
页数4
DOI
出版状态已出版 - 2009
活动2009 IEEE/WIC/ACM International Conference on Web Intelligence, WI 2009 - Milano, 意大利
期限: 15 9月 200918 9月 2009

出版系列

姓名Proceedings - 2009 IEEE/WIC/ACM International Conference on Web Intelligence, WI 2009
1

会议

会议2009 IEEE/WIC/ACM International Conference on Web Intelligence, WI 2009
国家/地区意大利
Milano
时期15/09/0918/09/09

学术指纹

探究 'Learning Deep Web crawling with diverse features' 的科研主题。它们共同构成独一无二的学术指纹。

引用此