跳到主要导航 跳到搜索 跳到主要内容

MTRFN: Multiscale Temporal Receptive Field Network for Compressed Video Action Recognition at Edge Servers

  • Xi'an Jiaotong University
  • Xinjiang University

科研成果: 期刊稿件文章同行评审

11 引用 (Scopus)

摘要

With the wide deployment of Internet of Things monitoring terminals, a tremendous number of videos are accumulated continuously. Big data processing and analysis-based action recognition has an increasingly important role in making cities simpler, better, and smarter. The traditional cloud server-centered analysis mode has to spend extra time transmitting vast video data terminals to remote cloud servers, which is always violated in real implementation. Edge servers with limited caching and computation capacities near the monitoring terminals enable implementation. However, due to the dependency on training data and the high complexity of extracting information and network architecture, existing image domain-based methods cannot be implemented at edge servers. Moreover, recognizing actions with different durations is still challenging. Due to these issues, we extend the traditional image domain to the compressed domain to efficiently extract the information of I frames and physical knowledge motion vectors (MVs), which can reflect the multiscale temporal feature just by partial decoding. To recognize the actions with different durations, a multiscale temporal receptive field network (MTRFN), including short-term and long-term branches, is proposed to simultaneously capture the action's instant change based on the extracted MVs, the long temporal feature between adjacent I frames, and the interaction between them. The results show that our algorithm can achieve a better balance between accuracy and computational complexity.

源语言英语
页(从-至)13965-13977
页数13
期刊IEEE Internet of Things Journal
9
15
DOI
出版状态已出版 - 1 8月 2022

学术指纹

探究 'MTRFN: Multiscale Temporal Receptive Field Network for Compressed Video Action Recognition at Edge Servers' 的科研主题。它们共同构成独一无二的指纹。

引用此