Skip to main navigation Skip to search Skip to main content

Action co-localization in an untrimmed video by graph neural networks

  • Xi'an Jiaotong University
  • HERE Global B.V.
  • Alibaba Group Holding Ltd.
  • Wormpex AI Research

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

3 Scopus citations

Abstract

We present an efficient approach for action co-localization in an untrimmed video by exploiting contextual and temporal feature from multiple action proposals. Most existing action localization methods focus on each individual action instances without accounting for the correlations among them. To exploit such correlations, we propose the Graph-based Temporal Action Co-Localization (G-TACL) method, which aggregates contextual features from multiple action proposals to assist temporal localization. This aggregation procedure is achieved with Graph Neural Networks with nodes initialized by the action proposal representations. In addition, a multi-level consistency evaluator is proposed to measure the similarity, which summarizes low-level temporal coincidences, features vector dot products and high-level contextual features similarities between any two proposals. Subsequently, these nodes are iteratively updated with Gated Recurrent Unit (GRU) and the obtained node features are used to regress the temporal boundaries of the action proposals, and finally to localize the action instances. Experiments on the THUMOS’14 and MEXaction2 datasets have demonstrated the efficacy of our proposed method.

Original languageEnglish
Title of host publicationMultiMedia Modeling - 26th International Conference, MMM 2020, Proceedings
EditorsWen-Huang Cheng, Junmo Kim, Jung-Woo Choi, Wei-Ta Chu, Peng Cui, Min-Chun Hu, Wesley De Neve
PublisherSpringer
Pages555-567
Number of pages13
ISBN (Print)9783030377304
DOIs
StatePublished - 2020
Event26th International Conference on MultiMedia Modeling, MMM 2020 - Daejeon, Korea, Republic of
Duration: 5 Jan 20208 Jan 2020

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume11961 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference26th International Conference on MultiMedia Modeling, MMM 2020
Country/TerritoryKorea, Republic of
CityDaejeon
Period5/01/208/01/20

Keywords

  • Multi-level consistency evaluator
  • Temporal action co-localization

Fingerprint

Dive into the research topics of 'Action co-localization in an untrimmed video by graph neural networks'. Together they form a unique fingerprint.

Cite this