跳到主要导航 跳到搜索 跳到主要内容

Momentum Centroid Alignment with Temporal-Relational Disentanglement for Cross-Domain Few-Shot Action Recognition

  • Xi'an Jiaotong University

科研成果: 期刊稿件文章同行评审

摘要

Traditional few-shot action recognition (FSAR) aims to address the problem of the scarcity of action videos, enabling the recognition of action categories with just a few labeled samples. It is generally believed that the samples in the metatraining phase and the meta-testing phase are all drawn from the same domain. However, in practical applications, they often come from different domains, which may lead to significant differences in the distribution of spatiotemporal features. Researchers have started to study the problem of cross-domain few-shot action recognition (CDFSAR). The current solution is to train the model by combining source domain video and unlabeled target domain video to improve the model’s generalization ability. In this paper, we follow this paradigm but make a more refined use of the unlabeled target domain videos to better extract transferable features. First, we decouple the source and target domain videos along the temporal dimension and extract the domain-irrelevant features in both the source and target domains. Second, in each episode, we calculate the centroid of the domain-irrelevant features of the target domain and perform a momentum update on this feature centroid. We use Cross-Attention to align the domain-irrelevant features of the source domain toward this dynamic centroid. Finally, we use these aligned source domain features for few-shot classification. Experimental results demonstrate that our approach significantly improves few-shot classification performance across diverse domain shifts, validating the effectiveness of our refined use of unlabeled target video.

学术指纹

探究 'Momentum Centroid Alignment with Temporal-Relational Disentanglement for Cross-Domain Few-Shot Action Recognition' 的科研主题。它们共同构成独一无二的学术指纹。

引用此