Skip to main navigation Skip to search Skip to main content

Spatio-temporal Collaborative Convolution for Video Action Recognition

  • Xi'an Jiaotong University
  • Peking University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Although video action recognition has achieved great progress in recent years, it is still a challenging task due to the huge computational complexity. Designing a lightweight network is a feasible solution, but it may reduce the spatio-temporal information modeling capability. In this paper, we propose a novel novel spatio-temporal collaborative convolution (denote as 'STC-Conv'), which can efficiently encode spatio-temporal information. STC-Conv collaboratively learn spatial and temporal feature in one convolution filter kernel. In short, temporal convolution and spatial convolution are integrated in the one STC convolution kernel, which can effectively reduce the model complexity and improve the computational efficiency. STC-Conv is a universal convolution, which can be applied to the existing 2D CNNs, such as ResNet, DenseNet. The experimental results on the temporal-related dataset Something Something V1 prove the superiority of our method. Noticeably, STC-Conv enjoys more excellent performance than 3D CNNs at even lower computation cost than standard 2D CNNs.

Original languageEnglish
Title of host publicationProceedings of 2020 IEEE International Conference on Artificial Intelligence and Computer Applications, ICAICA 2020
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages554-558
Number of pages5
ISBN (Electronic)9781728170046
DOIs
StatePublished - Jun 2020
Event2020 IEEE International Conference on Artificial Intelligence and Computer Applications, ICAICA 2020 - Dalian, China
Duration: 27 Jun 202029 Jun 2020

Publication series

NameProceedings of 2020 IEEE International Conference on Artificial Intelligence and Computer Applications, ICAICA 2020

Conference

Conference2020 IEEE International Conference on Artificial Intelligence and Computer Applications, ICAICA 2020
Country/TerritoryChina
CityDalian
Period27/06/2029/06/20

Keywords

  • Action Recognition
  • Spatio-Temporal Collaborative Convolution
  • Spatio-Temporal Modeling

Fingerprint

Dive into the research topics of 'Spatio-temporal Collaborative Convolution for Video Action Recognition'. Together they form a unique fingerprint.

Cite this