Skip to main navigation Skip to search Skip to main content

Correlation based file prefetching approach for Hadoop

  • Bo Dong
  • , Xiao Zhong
  • , Qinghua Zheng
  • , Lirong Jian
  • , Jian Liu
  • , Jie Qiu
  • , Ying Li

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

21 Scopus citations

Abstract

Hadoop Distributed File System (HDFS) has been widely adopted to support Internet applications because of its reliable, scalable and low-cost storage capability. BlueSky, one of the most popular e-Learning resource sharing systems in China, is utilizing HDFS to store massive courseware. However, due to the inefficient access mechanism of HDFS, access latency of reading files from HDFS significantly impacts the performance of processing user requests. This paper introduces a two-level correlation based file prefetching approach, taking the characteristics of HDFS into consideration, to improve performance by reducing access latency. Four placement patterns to store prefetched data are presented, with policies to achieve trade-off between performance and efficiency of HDFS prefetching. Moreover, a dynamic replica selection algorithm is investigated to improve the efficiency of HDFS prefetching. The proposed prefetching approach has been implemented in BlueSky, and experimental results prove that correlation based file prefetching can significantly reduce access latency therefore improve performance of Hadoop-based Internet applications.

Original languageEnglish
Title of host publicationProceedings - 2nd IEEE International Conference on Cloud Computing Technology and Science, CloudCom 2010
PublisherIEEE Computer Society
Pages41-48
Number of pages8
ISBN (Print)9780769543024
DOIs
StatePublished - 2010
Event2nd IEEE International Conference on Cloud Computing Technology and Science, CloudCom 2010 - Indianapolis, IN, United States
Duration: 30 Nov 20103 Dec 2010

Publication series

NameProceedings - 2nd IEEE International Conference on Cloud Computing Technology and Science, CloudCom 2010

Conference

Conference2nd IEEE International Conference on Cloud Computing Technology and Science, CloudCom 2010
Country/TerritoryUnited States
CityIndianapolis, IN
Period30/11/103/12/10

Keywords

  • Cloud storage
  • File correlation
  • Hadoop distributed file system
  • Prefetching

Fingerprint

Dive into the research topics of 'Correlation based file prefetching approach for Hadoop'. Together they form a unique fingerprint.

Cite this