TY - JOUR
T1 - MTRFN
T2 - Multiscale Temporal Receptive Field Network for Compressed Video Action Recognition at Edge Servers
AU - He, Lijun
AU - Zhang, Miao
AU - Zhang, Sijin
AU - Wang, Liejun
AU - Li, Fan
N1 - Publisher Copyright:
© 2014 IEEE.
PY - 2022/8/1
Y1 - 2022/8/1
N2 - With the wide deployment of Internet of Things monitoring terminals, a tremendous number of videos are accumulated continuously. Big data processing and analysis-based action recognition has an increasingly important role in making cities simpler, better, and smarter. The traditional cloud server-centered analysis mode has to spend extra time transmitting vast video data terminals to remote cloud servers, which is always violated in real implementation. Edge servers with limited caching and computation capacities near the monitoring terminals enable implementation. However, due to the dependency on training data and the high complexity of extracting information and network architecture, existing image domain-based methods cannot be implemented at edge servers. Moreover, recognizing actions with different durations is still challenging. Due to these issues, we extend the traditional image domain to the compressed domain to efficiently extract the information of I frames and physical knowledge motion vectors (MVs), which can reflect the multiscale temporal feature just by partial decoding. To recognize the actions with different durations, a multiscale temporal receptive field network (MTRFN), including short-term and long-term branches, is proposed to simultaneously capture the action's instant change based on the extracted MVs, the long temporal feature between adjacent I frames, and the interaction between them. The results show that our algorithm can achieve a better balance between accuracy and computational complexity.
AB - With the wide deployment of Internet of Things monitoring terminals, a tremendous number of videos are accumulated continuously. Big data processing and analysis-based action recognition has an increasingly important role in making cities simpler, better, and smarter. The traditional cloud server-centered analysis mode has to spend extra time transmitting vast video data terminals to remote cloud servers, which is always violated in real implementation. Edge servers with limited caching and computation capacities near the monitoring terminals enable implementation. However, due to the dependency on training data and the high complexity of extracting information and network architecture, existing image domain-based methods cannot be implemented at edge servers. Moreover, recognizing actions with different durations is still challenging. Due to these issues, we extend the traditional image domain to the compressed domain to efficiently extract the information of I frames and physical knowledge motion vectors (MVs), which can reflect the multiscale temporal feature just by partial decoding. To recognize the actions with different durations, a multiscale temporal receptive field network (MTRFN), including short-term and long-term branches, is proposed to simultaneously capture the action's instant change based on the extracted MVs, the long temporal feature between adjacent I frames, and the interaction between them. The results show that our algorithm can achieve a better balance between accuracy and computational complexity.
KW - Action recognition
KW - edge server
KW - multiscale
KW - physical knowledge
KW - video compressed domain
UR - https://www.scopus.com/pages/publications/85123302728
U2 - 10.1109/JIOT.2022.3142759
DO - 10.1109/JIOT.2022.3142759
M3 - 文章
AN - SCOPUS:85123302728
SN - 2327-4662
VL - 9
SP - 13965
EP - 13977
JO - IEEE Internet of Things Journal
JF - IEEE Internet of Things Journal
IS - 15
ER -