Skip to main navigation Skip to search Skip to main content

VRFormer: 360-Degree Video Streaming with FoV Combined Prediction and Super resolution

  • Zhihao Zhang
  • , Haipeng Du
  • , Shouqin Huang
  • , Weizhan Zhang
  • , Qinghua Zheng
  • Xi'an Jiaotong University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

5 Scopus citations

Abstract

360-degree video has shown great potential to the mainstream since its immersive experience. However, 360-degree video streaming requires ultrahigh bandwidth and low latency, which limit the improvement of user quality of experience (QoE). Currently, methods combining field of view (FoV) prediction and adaptive video streaming provide an effective method for addressing the above issues. However, existing FoV prediction methods based on recurrent neural networks (RNN) cannot capture long-range dependency from input to output. Current deep reinforcement learning (DRL)-based adaptive strategies fail to estimate the future bandwidth with high accuracy and fully explore the capability of VR devices. To ameliorate these limitations, we design a DRL-based 360-degree video streaming method named VRFormer with FoV combined prediction and super resolution (SR). First, we adopt a content-aware transformer-based encoder-decoder network to make the long-term FoV prediction. It combines the user's head movement history, eye-tracking history, and user attention extracted from a convolutional neural network (CNN)-based network. Second, we introduce a DNN-based SR network running on VR devices to reconstruct high-definition video content. Finally, we apply a DRL-based network to adaptively allocate rates for future tiles and dynamically control video content reconstruction. Experiments have verified that the proposed method can effectively improve the quality of experience (QoE) of the user's viewing experience compared to the state-of-the-art methods.

Original languageEnglish
Title of host publicationProceedings - 20th IEEE International Symposium on Parallel and Distributed Processing with Applications, 12th IEEE International Conference on Big Data and Cloud Computing, 12th IEEE International Conference on Sustainable Computing and Communications and 15th IEEE International Conference on Social Computing and Networking, ISPA/BDCloud/SocialCom/SustainCom 2022
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages531-538
Number of pages8
ISBN (Electronic)9781665464970
DOIs
StatePublished - 2022
Event20th IEEE International Symposium on Parallel and Distributed Processing with Applications, 12th IEEE International Conference on Big Data and Cloud Computing, 12th IEEE International Conference on Sustainable Computing and Communications and 15th IEEE International Conference on Social Computing and Networking, ISPA/BDCloud/SocialCom/SustainCom 2022 - Melbourne, Australia
Duration: 17 Dec 202219 Dec 2022

Publication series

NameProceedings - 20th IEEE International Symposium on Parallel and Distributed Processing with Applications, 12th IEEE International Conference on Big Data and Cloud Computing, 12th IEEE International Conference on Sustainable Computing and Communications and 15th IEEE International Conference on Social Computing and Networking, ISPA/BDCloud/SocialCom/SustainCom 2022

Conference

Conference20th IEEE International Symposium on Parallel and Distributed Processing with Applications, 12th IEEE International Conference on Big Data and Cloud Computing, 12th IEEE International Conference on Sustainable Computing and Communications and 15th IEEE International Conference on Social Computing and Networking, ISPA/BDCloud/SocialCom/SustainCom 2022
Country/TerritoryAustralia
CityMelbourne
Period17/12/2219/12/22

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 7 - Affordable and Clean Energy
    SDG 7 Affordable and Clean Energy

Keywords

  • 360-degree video streaming
  • DRL-based bit-rate adaptation
  • quality of experience
  • transformer-based FoV prediction

Fingerprint

Dive into the research topics of 'VRFormer: 360-Degree Video Streaming with FoV Combined Prediction and Super resolution'. Together they form a unique fingerprint.

Cite this