TY - GEN
T1 - Spatio-temporal Collaborative Convolution for Video Action Recognition
AU - Li, Xu
AU - Wen, Liqiang
AU - Wang, Jinjun
AU - Zeng, Ming
N1 - Publisher Copyright:
© 2020 IEEE.
PY - 2020/6
Y1 - 2020/6
N2 - Although video action recognition has achieved great progress in recent years, it is still a challenging task due to the huge computational complexity. Designing a lightweight network is a feasible solution, but it may reduce the spatio-temporal information modeling capability. In this paper, we propose a novel novel spatio-temporal collaborative convolution (denote as 'STC-Conv'), which can efficiently encode spatio-temporal information. STC-Conv collaboratively learn spatial and temporal feature in one convolution filter kernel. In short, temporal convolution and spatial convolution are integrated in the one STC convolution kernel, which can effectively reduce the model complexity and improve the computational efficiency. STC-Conv is a universal convolution, which can be applied to the existing 2D CNNs, such as ResNet, DenseNet. The experimental results on the temporal-related dataset Something Something V1 prove the superiority of our method. Noticeably, STC-Conv enjoys more excellent performance than 3D CNNs at even lower computation cost than standard 2D CNNs.
AB - Although video action recognition has achieved great progress in recent years, it is still a challenging task due to the huge computational complexity. Designing a lightweight network is a feasible solution, but it may reduce the spatio-temporal information modeling capability. In this paper, we propose a novel novel spatio-temporal collaborative convolution (denote as 'STC-Conv'), which can efficiently encode spatio-temporal information. STC-Conv collaboratively learn spatial and temporal feature in one convolution filter kernel. In short, temporal convolution and spatial convolution are integrated in the one STC convolution kernel, which can effectively reduce the model complexity and improve the computational efficiency. STC-Conv is a universal convolution, which can be applied to the existing 2D CNNs, such as ResNet, DenseNet. The experimental results on the temporal-related dataset Something Something V1 prove the superiority of our method. Noticeably, STC-Conv enjoys more excellent performance than 3D CNNs at even lower computation cost than standard 2D CNNs.
KW - Action Recognition
KW - Spatio-Temporal Collaborative Convolution
KW - Spatio-Temporal Modeling
UR - https://www.scopus.com/pages/publications/85092181454
U2 - 10.1109/ICAICA50127.2020.9182498
DO - 10.1109/ICAICA50127.2020.9182498
M3 - 会议稿件
AN - SCOPUS:85092181454
T3 - Proceedings of 2020 IEEE International Conference on Artificial Intelligence and Computer Applications, ICAICA 2020
SP - 554
EP - 558
BT - Proceedings of 2020 IEEE International Conference on Artificial Intelligence and Computer Applications, ICAICA 2020
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2020 IEEE International Conference on Artificial Intelligence and Computer Applications, ICAICA 2020
Y2 - 27 June 2020 through 29 June 2020
ER -