Skip to main navigation Skip to search Skip to main content

Momentum Centroid Alignment with Temporal-Relational Disentanglement for Cross-Domain Few-Shot Action Recognition

  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

Abstract

Traditional few-shot action recognition (FSAR) aims to address the problem of the scarcity of action videos, enabling the recognition of action categories with just a few labeled samples. It is generally believed that the samples in the metatraining phase and the meta-testing phase are all drawn from the same domain. However, in practical applications, they often come from different domains, which may lead to significant differences in the distribution of spatiotemporal features. Researchers have started to study the problem of cross-domain few-shot action recognition (CDFSAR). The current solution is to train the model by combining source domain video and unlabeled target domain video to improve the model’s generalization ability. In this paper, we follow this paradigm but make a more refined use of the unlabeled target domain videos to better extract transferable features. First, we decouple the source and target domain videos along the temporal dimension and extract the domain-irrelevant features in both the source and target domains. Second, in each episode, we calculate the centroid of the domain-irrelevant features of the target domain and perform a momentum update on this feature centroid. We use Cross-Attention to align the domain-irrelevant features of the source domain toward this dynamic centroid. Finally, we use these aligned source domain features for few-shot classification. Experimental results demonstrate that our approach significantly improves few-shot classification performance across diverse domain shifts, validating the effectiveness of our refined use of unlabeled target video.

Keywords

  • Cross-Domain
  • Few-shot Action Recognition
  • Momentum Centroid
  • Temporal-Relational Disentanglement

Fingerprint

Dive into the research topics of 'Momentum Centroid Alignment with Temporal-Relational Disentanglement for Cross-Domain Few-Shot Action Recognition'. Together they form a unique fingerprint.

Cite this