TY - GEN
T1 - Action co-localization in an untrimmed video by graph neural networks
AU - Zhai, Changbo
AU - Wang, Le
AU - Zhang, Qilin
AU - Gao, Zhanning
AU - Niu, Zhenxing
AU - Zheng, Nanning
AU - Hua, Gang
N1 - Publisher Copyright:
© 2020, Springer Nature Switzerland AG.
PY - 2020
Y1 - 2020
N2 - We present an efficient approach for action co-localization in an untrimmed video by exploiting contextual and temporal feature from multiple action proposals. Most existing action localization methods focus on each individual action instances without accounting for the correlations among them. To exploit such correlations, we propose the Graph-based Temporal Action Co-Localization (G-TACL) method, which aggregates contextual features from multiple action proposals to assist temporal localization. This aggregation procedure is achieved with Graph Neural Networks with nodes initialized by the action proposal representations. In addition, a multi-level consistency evaluator is proposed to measure the similarity, which summarizes low-level temporal coincidences, features vector dot products and high-level contextual features similarities between any two proposals. Subsequently, these nodes are iteratively updated with Gated Recurrent Unit (GRU) and the obtained node features are used to regress the temporal boundaries of the action proposals, and finally to localize the action instances. Experiments on the THUMOS’14 and MEXaction2 datasets have demonstrated the efficacy of our proposed method.
AB - We present an efficient approach for action co-localization in an untrimmed video by exploiting contextual and temporal feature from multiple action proposals. Most existing action localization methods focus on each individual action instances without accounting for the correlations among them. To exploit such correlations, we propose the Graph-based Temporal Action Co-Localization (G-TACL) method, which aggregates contextual features from multiple action proposals to assist temporal localization. This aggregation procedure is achieved with Graph Neural Networks with nodes initialized by the action proposal representations. In addition, a multi-level consistency evaluator is proposed to measure the similarity, which summarizes low-level temporal coincidences, features vector dot products and high-level contextual features similarities between any two proposals. Subsequently, these nodes are iteratively updated with Gated Recurrent Unit (GRU) and the obtained node features are used to regress the temporal boundaries of the action proposals, and finally to localize the action instances. Experiments on the THUMOS’14 and MEXaction2 datasets have demonstrated the efficacy of our proposed method.
KW - Multi-level consistency evaluator
KW - Temporal action co-localization
UR - https://www.scopus.com/pages/publications/85078535091
U2 - 10.1007/978-3-030-37731-1_45
DO - 10.1007/978-3-030-37731-1_45
M3 - 会议稿件
AN - SCOPUS:85078535091
SN - 9783030377304
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 555
EP - 567
BT - MultiMedia Modeling - 26th International Conference, MMM 2020, Proceedings
A2 - Cheng, Wen-Huang
A2 - Kim, Junmo
A2 - Choi, Jung-Woo
A2 - Chu, Wei-Ta
A2 - Cui, Peng
A2 - Hu, Min-Chun
A2 - De Neve, Wesley
PB - Springer
T2 - 26th International Conference on MultiMedia Modeling, MMM 2020
Y2 - 5 January 2020 through 8 January 2020
ER -